NaiveAI-Labs/Naive-N0.5-Flash
See how a 309 B MoE can be run with 15.5 B active parameters and why it matters for open‑source coding assistants.
The repo ships a 309‑billion‑parameter MoE transformer where each token routes to a small subset of experts, limiting the compute to roughly 15.5 billion active parameters. The architecture follows the Switch‑Transformer design: a shared backbone, a gating network, and dozens of feed‑forward experts that are sparsely activated. Training data is heavily weighted toward source code and related documentation, making the model a specialist for programming tasks and AI R&D code generation.
Inference hinges on the same sparse routing. By loading only the active experts for a batch, memory consumption drops to roughly 30 GB per GPU for a 2‑way expert activation, allowing a 4‑GPU server to run the model at acceptable latency. The gating logic is implemented in CUDA kernels that batch routing decisions, reducing per‑token overhead compared to dense models of similar scale. The repo provides scripts for sharding experts across GPUs and a simple API for token‑level generation.
Compared to open‑source dense models like LLaMA‑2‑70B, Naive‑N0.5‑Flash trades raw parameter count for capacity: the total parameter budget is an order of magnitude larger, but the active footprint is comparable to a 15‑B dense model. Early benchmarks on the HumanEval code‑completion suite show a 12‑percent improvement over LLaMA‑2‑70B, narrowing the gap to closed‑source Codex‑style systems. This demonstrates that MoE sparsity can deliver higher quality without proportional hardware cost.
The main trade‑off is routing latency and the need for synchronized expert execution. On a single GPU the model falls back to a dense mode, inflating memory to >120 GB and making it impractical. Scaling out adds network overhead, and the gating network can become a bottleneck under high request concurrency. Moreover, the expert selection is deterministic per token, which can amplify training data biases if certain experts dominate specific language patterns.
TakeawayNaive‑N0.5‑Flash proves a 309 B MoE can be made usable with only 15.5 B active parameters per token, unlocking near‑GPT‑4‑scale code models for the open‑source community.