🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Data

Min-K% Prob (Detecting Pretraining Data)

First page
Min-K% Prob (Detecting Pretraining Data)
Paper summary

Proposes Min-K% Prob as an effective detection method for determining whether specific text was in an LLM's pretraining data.

Ask this paper

Key points
01

Method: Computes the average log-probability of the K% least-likely tokens in a text; memorized text has higher log-probabilities on these tokens than unseen text.

02

Black-box detection: Works on API-accessible models without needing gradients or internal activations, making it broadly applicable.

03

Multiple use cases: Usable for benchmark-contamination detection, privacy auditing of machine unlearning, and copyrighted-text detection in pretraining corpora.

04

Policy implications: Provides a technical tool for the copyright and privacy debates, letting third parties measurably test specific-text inclusion in training data.

Every Monday
Get next week’s papers.
Subscribe on Substack