🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency · Architecture · Training

Puzzle-75B

First page
Puzzle-75B
Paper summary

Bigger mixture-of-experts models keep winning on quality, but serving them at interactive latency is still hard. NVIDIA compresses the hybrid MoE Nemotron-3-Super into Puzzle-75B-A9B and roughly doubles interactive server throughput while holding quality.

Ask this paper

Key points
01

Joint structural search: Heterogeneous MoE pruning, active-parameter budget, and Mamba pruning are optimized together rather than one at a time, wrapped in an iterative pipeline with distillation, RL, quantization, and a Multi-Token Prediction head.

02

Large throughput gains: On a single 8xB200 node it hits about 2x the parent's server throughput at matched user-throughput, a direct win for anyone serving these models under latency constraints.

03

Concurrency at long context: At 1M-token context on a single H100, concurrency climbs from 1 request to 8, expanding what long-context workloads a single accelerator can host.

04

Why it matters: Accuracy holds across reasoning, coding, long-context, and agentic benchmarks, so cheaper serving with agentic capability intact changes what teams can afford to run in production.

Every Monday
Get next week’s papers.
Subscribe on Substack