
Llemma
Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

LLMs for Software Engineering
A comprehensive survey of LLMs for software engineering covering models, tasks, evaluation, and open challenges.

Self-RAG
Self-RAG trains an LM to adaptively retrieve, generate, and self-critique using special reflection tokens.

RAG for Long-Form QA
Explores retrieval-augmented LMs specifically on long-form question answering, where RAG failures are more subtle.

GenBench
A Nature Machine Intelligence paper framework for characterizing and understanding generalization research in NLP.

LLM Self-Explanations
Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

OpenAgents
An open platform for running and hosting real-world language agents, including three distinct agent types.

Eliciting Human Preferences with LLMs
Anthropic uses LLMs to guide the task-specification process, eliciting user intent through natural-language dialogue.

AutoMix
AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.

Video Language Planning
Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack