Forecasting Scientific Progress with AI
Free while signed in. Answers cite the passages they came from.

Can frontier models predict where science is going? This work introduces CUSP, a cutoff-conditioned benchmark built from 4,760 real scientific events across multiple disciplines, each grounded against a verified knowledge cutoff. For every event, models are tested on four tasks: feasibility assessment, mechanistic reasoning, generative solution design, and temporal prediction. The headline is sobering: models recognize plausible directions but cannot forecast outcomes.
Recognition is not foresight: Models can identify plausible research directions when choosing among competing candidates, but they fail to reliably predict whether an advance will actually be realized, and they systematically misestimate when it will happen.
Domain-dependent, and timing is hardest: Performance is highly heterogeneous across fields, with the timing of AI progress more predictable than advances in biology, chemistry, and physics. Temporal prediction is the weakest skill across the board.
Not just a training-cutoff artifact: Performance is largely insensitive to whether an event falls before or after the model's training cutoff. Extra pre-cutoff knowledge helps but does not close the gap to full-information settings, and that gap widens for high-citation advances.
Why it matters: Models also show systematic overconfidence and strong response biases, which means unreliable uncertainty estimates. As labs lean on AI to triage research bets, CUSP gives a controlled way to measure where it helps, surfacing directions, and where it fails, predicting outcomes.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack