🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture

Arctic

Paper preview
Arctic
Paper summary

Snowflake's Arctic is an Apache 2.0 open LLM with a Dense-MoE Hybrid transformer (480B total / 17B active) that matches Llama 3 70B on enterprise metrics while using under 3K GPU weeks (~$2M) of training compute - roughly 17x less than Llama 3 70B.

Ask this paper

Key points
01

Dense-MoE Hybrid architecture: A 10B dense backbone is paired with 128 fine-grained 3.66B experts selected via top-2 gating, balancing model capacity with communication efficiency across GPUs.

02

Enterprise-first curriculum: A three-stage data curriculum emphasizes SQL, coding, and instruction-following over broad world knowledge, yielding strong results on Spider (SQL), HumanEval+/MBPP+ (code), and IFEval.

03

Compute efficiency: Training ran in under 3K GPU weeks for roughly $2M, enabled by architecture-system co-design that overlaps communication with computation to hide MoE routing latency.

04

Fully open release: Apache 2.0 weights, training code, serving code, fine-tuning pipelines, and a detailed cookbook, available across HuggingFace, NVIDIA API, AWS, and Azure.

Every Monday
Get next week’s papers.
Subscribe on Substack