Puzzle-75B
Free while signed in. Answers cite the passages they came from.

Bigger mixture-of-experts models keep winning on quality, but serving them at interactive latency is still hard. NVIDIA compresses the hybrid MoE Nemotron-3-Super into Puzzle-75B-A9B and roughly doubles interactive server throughput while holding quality.
Joint structural search: Heterogeneous MoE pruning, active-parameter budget, and Mamba pruning are optimized together rather than one at a time, wrapped in an iterative pipeline with distillation, RL, quantization, and a Multi-Token Prediction head.
Large throughput gains: On a single 8xB200 node it hits about 2x the parent's server throughput at matched user-throughput, a direct win for anyone serving these models under latency constraints.
Concurrency at long context: At 1M-token context on a single H100, concurrency climbs from 1 request to 8, expanding what long-context workloads a single accelerator can host.
Why it matters: Accuracy holds across reasoning, coding, long-context, and agentic benchmarks, so cheaper serving with agentic capability intact changes what teams can afford to run in production.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack