🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning

Advancing LLM Reasoning (Eurus)

Free while signed in. Answers cite the passages they came from.

First page
Advancing LLM Reasoning (Eurus)
The curator’s take

OpenBMB's Eurus is a suite of reasoning-specialized LLMs (7B and 70B) fine-tuned on UltraInteract, a new alignment dataset built around preference trees for complex math, code, and logical tasks.

Key points
01

UltraInteract data: Each instruction is paired with a tree of reasoning chains plus multi-turn interactions and pairwise preferences, giving the model structured examples of correct vs incorrect reasoning.

02

SoTA open reasoning: Across 12 reasoning benchmarks Eurus-70B surpasses GPT-3.5 Turbo, reaching 33.3% on LeetCode and 32.6% on TheoremQA, and beats existing open-source baselines by 13.3%+ on average.

03

Tailored reward objective: Standard DPO is shown to be suboptimal for reasoning, so the authors design a specialized reward modeling objective better suited to chain-of-thought style data.

04

Takeaway: Demonstrates that task-specific alignment data design - not just model scale - is a key lever for pushing open models into frontier reasoning territory.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack