🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Reasoning

Evaluation of o1

First page
Evaluation of o1
Paper summary

provides a comprehensive evaluation of OpenAI's o1-preview LLM; shows strong performance across many tasks such as competitive programming, generating coherent and accurate radiology reports, high school-level mathematical reasoning tasks, chip design tasks, anthropology and geology, quantitative investing, social media analysis, and many other domains and problems.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack