Min-K% Prob (Detecting Pretraining Data)
Free while signed in. Answers cite the passages they came from.

Proposes Min-K% Prob as an effective detection method for determining whether specific text was in an LLM's pretraining data.
Method: Computes the average log-probability of the K% least-likely tokens in a text; memorized text has higher log-probabilities on these tokens than unseen text.
Black-box detection: Works on API-accessible models without needing gradients or internal activations, making it broadly applicable.
Multiple use cases: Usable for benchmark-contamination detection, privacy auditing of machine unlearning, and copyrighted-text detection in pretraining corpora.
Policy implications: Provides a technical tool for the copyright and privacy debates, letting third parties measurably test specific-text inclusion in training data.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack