🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

Advancing LLM Reasoning (Eurus)

First page
Advancing LLM Reasoning (Eurus)
Paper summary

OpenBMB's Eurus is a suite of reasoning-specialized LLMs (7B and 70B) fine-tuned on UltraInteract, a new alignment dataset built around preference trees for complex math, code, and logical tasks.

Ask this paper

Key points
01

UltraInteract data: Each instruction is paired with a tree of reasoning chains plus multi-turn interactions and pairwise preferences, giving the model structured examples of correct vs incorrect reasoning.

02

SoTA open reasoning: Across 12 reasoning benchmarks Eurus-70B surpasses GPT-3.5 Turbo, reaching 33.3% on LeetCode and 32.6% on TheoremQA, and beats existing open-source baselines by 13.3%+ on average.

03

Tailored reward objective: Standard DPO is shown to be suboptimal for reasoning, so the authors design a specialized reward modeling objective better suited to chain-of-thought style data.

04

Takeaway: Demonstrates that task-specific alignment data design - not just model scale - is a key lever for pushing open models into frontier reasoning territory.

Every Monday
Get next week’s papers.
Subscribe on Substack