🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning

DeepSeekMath

Free while signed in. Answers cite the passages they came from.

First page
DeepSeekMath
The curator’s take

DeepSeek releases DeepSeekMath 7B, a math-specialized LLM that closes much of the gap to GPT-4 and Gemini-Ultra on MATH by combining better data and a new RL objective.

Key points
01

120B math pretraining: Continues pretraining a DeepSeek-Coder base on 120B math-related tokens mined from Common Crawl, plus natural language and code, so that math-relevant skills and knowledge sit at the foundation.

02

GRPO objective: Introduces Group Relative Policy Optimization, a PPO variant that drops the value network and estimates advantage via within-group comparisons, reducing memory with strong reasoning results.

03

Close to frontier: DeepSeekMath 7B reaches 51.7% on MATH, approaching Gemini-Ultra (53.2%) and GPT-4 (52.9%) without any external tools.

04

Self-consistency boost: Combining DeepSeekMath 7B with self-consistency over 64 samples pushes performance to 60.9%, beating frontier API models on MATH.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack