How Code Empowers LLMs
Free while signed in. Answers cite the passages they came from.

A survey on why training LLMs with code data produces capabilities well beyond coding itself.
Capability catalog: Training on code improves code generation, general reasoning, structured-output fidelity, function calling, and agentic behavior - a single intervention with broad downstream effects.
Reasoning link: Evidence that code data strengthens step-by-step reasoning even on non-code tasks, supporting the argument that programming languages provide cleaner long-range dependency signals than natural language.
Tool use and agents: Code pretraining is positioned as a prerequisite for tool-calling and agent behaviors, where models must emit structured, executable artifacts.
Open directions: Identifies gaps in understanding *which* code properties matter most (syntax discipline, execution traces, type structure) - still a live research question.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack