🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Measuring Higher Level Mathematical Reasoning

Full-paper indexing in progress
Paper preview
Measuring Higher Level Mathematical Reasoning
Paper summary

introduces Putnam-AXIOM, a new math reasoning benchmark with 236 Putnam Competition problems and 52 variations; even the best model considered (OpenAI's o1-preview) achieves only 41.95% accuracy on original problems and performs significantly worse on variations.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack