Survey on Language Models for Code
First page

Paper summary
A comprehensive survey of LLMs for code covering 50+ models, 30+ evaluation tasks, and 500 related works.
Ask this paper
01
Model landscape: Catalogs 50+ code LLMs across sizes, architectures, and training regimes, providing a single reference for what's available.
02
Task taxonomy: Reviews 30+ evaluation tasks spanning code generation, repair, translation, summarization, and execution prediction.
03
Training and data recipes: Walks through pretraining corpus construction, instruction tuning, and RLHF specifically for code.
04
Open problems: Highlights challenges in long-context code understanding, multi-file reasoning, and robust evaluation beyond HumanEval-style metrics.