Advancing LLM Reasoning (Eurus)
Free while signed in. Answers cite the passages they came from.

OpenBMB's Eurus is a suite of reasoning-specialized LLMs (7B and 70B) fine-tuned on UltraInteract, a new alignment dataset built around preference trees for complex math, code, and logical tasks.
UltraInteract data: Each instruction is paired with a tree of reasoning chains plus multi-turn interactions and pairwise preferences, giving the model structured examples of correct vs incorrect reasoning.
SoTA open reasoning: Across 12 reasoning benchmarks Eurus-70B surpasses GPT-3.5 Turbo, reaching 33.3% on LeetCode and 32.6% on TheoremQA, and beats existing open-source baselines by 13.3%+ on average.
Tailored reward objective: Standard DPO is shown to be suboptimal for reasoning, so the authors design a specialized reward modeling objective better suited to chain-of-thought style data.
Takeaway: Demonstrates that task-specific alignment data design - not just model scale - is a key lever for pushing open models into frontier reasoning territory.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack