🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Data

MLE-Bench

First page
MLE-Bench
Paper summary

proposes a new benchmark for the evaluation of machine learning agents on machine learning engineering capabilities; includes 75 ML engineering-related competition from Kaggle testing on MLE skills such as training models, preparing datasets, and running experiments; OpenAI’s o1-preview with the AIDE scaffolding achieves Kaggle bronze medal level in 16.9% of competitions.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack