🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Reliability of Watermarks for LLMs

First page
Reliability of Watermarks for LLMs
Paper summary

Studies whether watermarks survive human rewriting and LLM paraphrasing.

Ask this paper

Key points
01

Robustness testing: Evaluates whether watermarks remain detectable after human rewrites, paraphrasing attacks, and translation round-trips.

02

Surprisingly robust: Finds that statistical watermarks (Kirchenbauer et al.) remain detectable even after aggressive transformations, with enough output text.

03

Text-length dependence: Detection confidence scales with text length - short watermarked snippets are much easier to obliterate than long ones.

04

AI detection realism: Provides a sober evaluation of watermarking's practical viability amid concerns about AI-generated content.

Every Monday
Get next week’s papers.
Subscribe on Substack