Puzzle-75B

Bigger mixture-of-experts models keep winning on quality, but serving them at interactive latency is still hard. NVIDIA compresses the hybrid MoE Nemotron-3-Super into Puzzle-75B-A9B and roughly doubles interactive server throughput while holding quality.
Ask this paper
Joint structural search: Heterogeneous MoE pruning, active-parameter budget, and Mamba pruning are optimized together rather than one at a time, wrapped in an iterative pipeline with distillation, RL, quantization, and a Multi-Token Prediction head.
Large throughput gains: On a single 8xB200 node it hits about 2x the parent's server throughput at matched user-throughput, a direct win for anyone serving these models under latency constraints.
Concurrency at long context: At 1M-token context on a single H100, concurrency climbs from 1 request to 8, expanding what long-context workloads a single accelerator can host.
Why it matters: Accuracy holds across reasoning, coding, long-context, and agentic benchmarks, so cheaper serving with agentic capability intact changes what teams can afford to run in production.