🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Safety

SimpleQA

Paper preview
SimpleQA
Paper summary

a challenging benchmark of 4,326 short factual questions adversarially collected against GPT-4 responses; reports that frontier models like GPT-4o and Claude achieve less than 50% accuracy; finds that there is a positive calibration between the model stated confidence and accuracy, signaling that they have some notion of confidence; claims that there is still room to improve the calibration of LLMs in terms of stated confidence.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack