🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Code · Efficiency

JAXBench

Free while signed in. Answers cite the passages they came from.

First page
JAXBench
The curator’s take

GPU kernel optimization has KernelBench to hillclimb on. TPUs had nothing, and the Pallas DSL is documented thinly enough that models mostly guess. Google, with Harvard and UC Berkeley, closes that gap and finds a clean lesson about context along the way.

Key points
01

Built from production workloads: JAXBench holds 50 JAX workloads, 17 production ML operators extracted from MaxText architectures such as Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, and AlphaFold2, plus 33 operators translated from KernelBench and resized for high TPU v6e MXU utilization.

02

Measured against experts: Eight of the 17 production operators ship with hand-optimized Pallas kernels from the public Tokamax library, block-size tuned, so agent output gets compared to expert work instead of a naive baseline.

03

Context beats scale: With Gemini 3 Flash, conditioning on curated TPU documentation raises per-sample correctness from 5.8% to 37.3% and solves 48 of 50 benchmarks at a 1.28x geomean speedup, while Autocomp's beam search pushes it to 1.36x and reaches 1.60x on the hand-tuned subset.

04

Why it matters: Correctness turned out to be a documentation problem and speed turned out to be a search problem, a split that generalizes to any agent working against an API it was never trained on.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack