Advancing LLM Reasoning (Eurus)

OpenBMB's Eurus is a suite of reasoning-specialized LLMs (7B and 70B) fine-tuned on UltraInteract, a new alignment dataset built around preference trees for complex math, code, and logical tasks.
Ask this paper
UltraInteract data: Each instruction is paired with a tree of reasoning chains plus multi-turn interactions and pairwise preferences, giving the model structured examples of correct vs incorrect reasoning.
SoTA open reasoning: Across 12 reasoning benchmarks Eurus-70B surpasses GPT-3.5 Turbo, reaching 33.3% on LeetCode and 32.6% on TheoremQA, and beats existing open-source baselines by 13.3%+ on average.
Tailored reward objective: Standard DPO is shown to be suboptimal for reasoning, so the authors design a specialized reward modeling objective better suited to chain-of-thought style data.
Takeaway: Demonstrates that task-specific alignment data design - not just model scale - is a key lever for pushing open models into frontier reasoning territory.