🚀NEW LABGetting Started with Claude AgentsStart lab
Data · Training · Reasoning

Llemma

First page
Llemma
Paper summary

Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

Ask this paper

Key points
01

Proof-Pile-2 dataset: Mixes scientific papers, math-heavy web pages, and mathematical code into a focused math-pretraining corpus.

02

Code Llama base: Uses Code Llama as the base model, leveraging its existing code proficiency as a scaffold for formal-style math reasoning.

03

Beats unreleased Minerva: Outperforms open base models and the unreleased Minerva on the MATH benchmark at comparable scale.

04

Full open release: Releases model, dataset, and code - positioning Llemma as a reproducible starting point for open mathematical LLM research.

Every Monday
Get next week’s papers.
Subscribe on Substack