LM In-Context Recall is Prompt Dependent

Using needle-in-a-haystack tests across multiple models, this paper shows that in-context recall is highly sensitive to prompt wording and that training data biases can silently degrade a model's ability to retrieve from its own context.
Ask this paper
Needle-in-a-haystack methodology: A factoid is embedded at various positions inside a long filler context, and recall is measured as prompt length and needle depth vary across several frontier LLMs.
Prompt sensitivity: Small rewordings of the query can move recall accuracy dramatically, indicating that existing "context-window size" numbers overstate practical recall capability.
Training-data interference: When the needle conflicts with content the model likely saw in pre-training, the model often returns its memorized answer instead of the in-context fact - a subtle but important failure mode.
Paths to improve recall: The paper shows that larger size, stronger attention mechanisms, alternative training objectives, and targeted fine-tuning each independently improve recall under prompt variation.