🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

Many-shot Jailbreaking

Free while signed in. Answers cite the passages they came from.

Paper preview
Many-shot Jailbreaking
The curator’s take

Anthropic shows that long-context windows enable a new attack where hundreds of fake user/assistant dialogues are packed into a single prompt, coaxing frontier LLMs to answer the final harmful question despite safety training.

Key points
01

Attack mechanics: The prompt embeds faux dialogues in which a cooperating assistant answers harmful queries, followed by the target question; in-context learning generalizes the pattern and overrides RLHF safety.

02

Power-law scaling: Attack success rises as a power law in the number of shots up to ~256, and is more effective on larger models that have stronger in-context learning.

03

Model coverage: The technique works on Claude 2.0 and comparable frontier LLMs from other labs, indicating a structural rather than model-specific vulnerability.

04

Mitigations: Context-window limits plus prompt classification / modification cut attack success from 61% to 2% in one setup, though the authors note variants may evade detection.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack