🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

Reliability of Watermarks for LLMs

Free while signed in. Answers cite the passages they came from.

First page
Reliability of Watermarks for LLMs
The curator’s take

Studies whether watermarks survive human rewriting and LLM paraphrasing.

Key points
01

Robustness testing: Evaluates whether watermarks remain detectable after human rewrites, paraphrasing attacks, and translation round-trips.

02

Surprisingly robust: Finds that statistical watermarks (Kirchenbauer et al.) remain detectable even after aggressive transformations, with enough output text.

03

Text-length dependence: Detection confidence scales with text length - short watermarked snippets are much easier to obliterate than long ones.

04

AI detection realism: Provides a sober evaluation of watermarking's practical viability amid concerns about AI-generated content.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack