Arctic
Free while signed in. Answers cite the passages they came from.

Snowflake's Arctic is an Apache 2.0 open LLM with a Dense-MoE Hybrid transformer (480B total / 17B active) that matches Llama 3 70B on enterprise metrics while using under 3K GPU weeks (~$2M) of training compute - roughly 17x less than Llama 3 70B.
Dense-MoE Hybrid architecture: A 10B dense backbone is paired with 128 fine-grained 3.66B experts selected via top-2 gating, balancing model capacity with communication efficiency across GPUs.
Enterprise-first curriculum: A three-stage data curriculum emphasizes SQL, coding, and instruction-following over broad world knowledge, yielding strong results on Spider (SQL), HumanEval+/MBPP+ (code), and IFEval.
Compute efficiency: Training ran in under 3K GPU weeks for roughly $2M, enabled by architecture-system co-design that overlaps communication with computation to hide MoE routing latency.
Fully open release: Apache 2.0 weights, training code, serving code, fine-tuning pipelines, and a detailed cookbook, available across HuggingFace, NVIDIA API, AWS, and Azure.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack