LM In-Context Recall is Prompt Dependent
Free while signed in. Answers cite the passages they came from.

Using needle-in-a-haystack tests across multiple models, this paper shows that in-context recall is highly sensitive to prompt wording and that training data biases can silently degrade a model's ability to retrieve from its own context.
Needle-in-a-haystack methodology: A factoid is embedded at various positions inside a long filler context, and recall is measured as prompt length and needle depth vary across several frontier LLMs.
Prompt sensitivity: Small rewordings of the query can move recall accuracy dramatically, indicating that existing "context-window size" numbers overstate practical recall capability.
Training-data interference: When the needle conflicts with content the model likely saw in pre-training, the model often returns its memorized answer instead of the in-context fact - a subtle but important failure mode.
Paths to improve recall: The paper shows that larger size, stronger attention mechanisms, alternative training objectives, and targeted fine-tuning each independently improve recall under prompt variation.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack