Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification

Joshua Ong Jun Leang and colleagues (MBZUAI, Imperial, Edinburgh and UCL) build Magenta, a training-free agent that answers a math problem, restates the answer in Lean 4 and produces a machine-checked proof, reaching 100% on recent olympiad benchmarks.
Ask this paper
Pipeline: From a natural-language problem the agent produces an answer, formalizes it as a Lean 4 statement and constructs a proof that Lean checks.
Statement judge: A judge checks that the formal statement preserves the original problem, which prevents proofs of a different, easier statement.
Error attribution: A second judge sends failed attempts either back to mathematical re-derivation or to local Lean repair.
Results: 100% on all evaluated olympiad benchmarks, including AIME 2025, AIME 2026 and HMMT February 2026. With the open-weight K2-Horizon-7B reasoner it solves all six IMO 2026 problems.
Ablations: Statement adjudication is necessary to avoid false certificates, and feedback-guided correction beats independent resampling on hard problems.
Abstract
Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathematical reasoning. Restraining LLMs to informal reasoning misses out on the opportunity to use the discrete verification abilities that machines offer through machine-checkable proofs. In this paper, we bridge the gap between informal and formal reasoning by integrating Lean signals into the informal reasoning process. We introduce Magenta, a training-free agentic pipeline that, given only a natural-language problem, produces an answer, expresses it as a Lean 4 statement, and constructs a machine-checked proof. A statement judge verifies whether the formalisation preserves the original problem, while an error-attribution judge routes failed attempts either to mathematical re-derivation or local Lean repair. Magenta achieves 100% accuracy across all evaluated olympiad benchmarks, including AIME 2025, AIME 2026, and HMMT February 2026. When paired with the open-weight K2-Horizon-7B reasoner, it solves all six IMO 2026 problems. Our analysis shows that statement adjudication is essential for preventing false certificates and that feedback-guided correction outperforms independent resampling on difficult problems.