OLMOTrace
First page

Paper summary
Allen Institute for AI & University of Washington present OLMOTRACE, a real-time system that traces LLM-generated text back to its verbatim sources in the original training data, even across multi-trillion-token corpora.
Ask this paper
01
What it does: For a given LM output, OLMOTRACE highlights exact matches with training data segments and lets users inspect full documents for those matches. Think "reverse-engineering" a model’s response via lexical lookup.
02
How it works:
03
Supported models: Works with OLMo models (e.g., OLMo-2-32B-Instruct) and their full pre/mid/post-training datasets, totaling 4.6T tokens.
04
Use cases:
05
Benchmarked:
06
Not RAG: It retrieves after generation, without changing output, unlike retrieval-augmented generation.