Survey on Language Models for Code
Free while signed in. Answers cite the passages they came from.

A comprehensive survey of LLMs for code covering 50+ models, 30+ evaluation tasks, and 500 related works.
Model landscape: Catalogs 50+ code LLMs across sizes, architectures, and training regimes, providing a single reference for what's available.
Task taxonomy: Reviews 30+ evaluation tasks spanning code generation, repair, translation, summarization, and execution prediction.
Training and data recipes: Walks through pretraining corpus construction, instruction tuning, and RLHF specifically for code.
Open problems: Highlights challenges in long-context code understanding, multi-file reasoning, and robust evaluation beyond HumanEval-style metrics.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack