🚀NEW LABGetting Started with Claude AgentsStart lab
Code · Data

Naturalized Execution Tuning (NExT)

First page
Naturalized Execution Tuning (NExT)
Paper summary

NExT teaches LLMs to reason about program runtime behavior by generating synthetic chain-of-thought rationales over execution traces. The approach bootstraps training data through self-training rather than manual annotation, and the learned reasoning transfers to scenarios where traces are unavailable at inference.

Ask this paper

Key points
01

Execution-aware rationales: The model inspects variable states at each step of a program's execution and produces natural-language rationales explaining what is happening, then fine-tunes on those rationales.

02

Self-training bootstrap: No human-annotated rationales are needed - rationales for correct repairs are kept and used to teach the model, scaling naturally across many programs.

03

Large repair gains: On PaLM 2, NExT improves fix rate by +26.1% on MBPP and +14.3% on HumanEval (absolute), with both automated metrics and human evaluators rating the rationales as higher quality.

04

Trace-free generalization: At inference the model applies the same reasoning patterns without live execution traces, showing it learned transferable execution-aware reasoning rather than trace-copying.

Every Monday
Get next week’s papers.
Subscribe on Substack