How Code Empowers LLMs

A survey on why training LLMs with code data produces capabilities well beyond coding itself.
Ask this paper
Capability catalog: Training on code improves code generation, general reasoning, structured-output fidelity, function calling, and agentic behavior - a single intervention with broad downstream effects.
Reasoning link: Evidence that code data strengthens step-by-step reasoning even on non-code tasks, supporting the argument that programming languages provide cleaner long-range dependency signals than natural language.
Tool use and agents: Code pretraining is positioned as a prerequisite for tool-calling and agent behaviors, where models must emit structured, executable artifacts.
Open directions: Identifies gaps in understanding *which* code properties matter most (syntax discipline, execution traces, type structure) - still a live research question.