LLMs Predict Neuroscience Results (BrainBench)
Free while signed in. Answers cite the passages they came from.

BrainBench asks both LLMs and human experts to predict the outcomes of neuroscience experiments from their abstracts, and finds LLMs outperform experts.
Benchmark design: Each item is a real neuroscience abstract with two alternative result endings (correct and plausibly incorrect); the task is to pick the correct one.
LLMs beat experts: Across several frontier LLMs, mean accuracy is higher than the mean of human neuroscientists, flipping the usual expectation that experts dominate specialist benchmarks.
BrainGPT: A version fine-tuned on neuroscience literature (BrainGPT) improves further, showing that targeted pretraining can extract more structure from a specialist corpus.
Calibrated confidence: Like human experts, LLMs are more accurate when they express higher confidence, pointing toward productive human-AI collaboration in forward-looking science.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack