Arctic

Snowflake's Arctic is an Apache 2.0 open LLM with a Dense-MoE Hybrid transformer (480B total / 17B active) that matches Llama 3 70B on enterprise metrics while using under 3K GPU weeks (~$2M) of training compute - roughly 17x less than Llama 3 70B.
Ask this paper
Dense-MoE Hybrid architecture: A 10B dense backbone is paired with 128 fine-grained 3.66B experts selected via top-2 gating, balancing model capacity with communication efficiency across GPUs.
Enterprise-first curriculum: A three-stage data curriculum emphasizes SQL, coding, and instruction-following over broad world knowledge, yielding strong results on Spider (SQL), HumanEval+/MBPP+ (code), and IFEval.
Compute efficiency: Training ran in under 3K GPU weeks for roughly $2M, enabled by architecture-system co-design that overlaps communication with computation to hide MoE routing latency.
Fully open release: Apache 2.0 weights, training code, serving code, fine-tuning pipelines, and a detailed cookbook, available across HuggingFace, NVIDIA API, AWS, and Azure.