How Faithful are RAG Models? (ClashEval)
Free while signed in. Answers cite the passages they came from.

ClashEval constructs a 1,200-question benchmark across six domains with intentionally corrupted retrieved documents to measure when RAG helps and when it misleads GPT-4 and other top LLMs.
Controlled conflict setup: Questions span drug dosages, Olympic records, locations, and other verifiable facts. Retrieved documents are perturbed from subtle to obvious errors so the authors can study model behavior under realistic vs implausible corruption.
RAG can override correct priors: LLMs abandon their correct internal knowledge over 60% of the time when the retrieved document is wrong but plausible - a strong demonstration that retrieval can hurt as much as help.
Prior strength matters: The weaker a model's initial confidence (measured via token probabilities), the more it capitulates to the retrieved content. Models with strong, high-probability priors resist incorrect retrieval more effectively.
Simple interventions help: The authors show that exploiting confidence signals - for example, gating acceptance of retrieved content on the model's prior probability - measurably improves accuracy under conflicting information.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack