Reliability of Watermarks for LLMs
Free while signed in. Answers cite the passages they came from.

Studies whether watermarks survive human rewriting and LLM paraphrasing.
Robustness testing: Evaluates whether watermarks remain detectable after human rewrites, paraphrasing attacks, and translation round-trips.
Surprisingly robust: Finds that statistical watermarks (Kirchenbauer et al.) remain detectable even after aggressive transformations, with enough output text.
Text-length dependence: Detection confidence scales with text length - short watermarked snippets are much easier to obliterate than long ones.
AI detection realism: Provides a sober evaluation of watermarking's practical viability amid concerns about AI-generated content.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack