🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 4, 2026
Retrieval · Evaluation

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

First page
Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
The curator’s take

Matthias Busch and colleagues at Helmholtz-Zentrum Hereon and Hamburg University of Technology audit 22 frontier models on 12 molecular regression benchmarks to separate models that predict a property from models that reproduce a published number.

Ask this paper

Key points
01

Prevalence: Verbatim retrieval is widespread but benchmark-specific. On five of the twelve datasets more than 50% of the models show it; on the rest it appears only in isolated cells.

02

Reasoning makes it worse: The same experiments, same molecules, same prompt, are flagged 89% more often at the higher reasoning level than at the lowest one.

03

Interrupting retrieval: Transforming SMILES strings does not fully block it. The strongest models sometimes still recognize a transformed SMILES paired with an original label.

04

What suppression reveals: Suppressing retrieval pulls the models' relative prediction errors closer together, while differing use of verbatim retrieval spreads them apart. Predictive capability therefore accounts for part of the ranking on these benchmarks and memorized values account for another part.

Abstract

Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than $50\%$ of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged $89\%$ more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cases, and find that the strongest models in some cases still recognise a combination of transformed SMILES strings and original labels. Furthermore, suppressing retrieval moves the prediction errors of the different models closer together in relative terms, while their differing use of verbatim retrieval spreads them apart. This indicates that the general predictive capability of an LLM is not determined solely by the amount of memorised values. This work provides an overview of the amount and depth of verbatim retrieval in molecular regression benchmarks using LLMs.

Every Monday
Get next week’s papers.
Subscribe on Substack