Reliability of Watermarks for LLMs
First page

Paper summary
Studies whether watermarks survive human rewriting and LLM paraphrasing.
Ask this paper
01
Robustness testing: Evaluates whether watermarks remain detectable after human rewrites, paraphrasing attacks, and translation round-trips.
02
Surprisingly robust: Finds that statistical watermarks (Kirchenbauer et al.) remain detectable even after aggressive transformations, with enough output text.
03
Text-length dependence: Detection confidence scales with text length - short watermarked snippets are much easier to obliterate than long ones.
04
AI detection realism: Provides a sober evaluation of watermarking's practical viability amid concerns about AI-generated content.