🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Many-shot Jailbreaking

Paper preview
Many-shot Jailbreaking
Paper summary

Anthropic shows that long-context windows enable a new attack where hundreds of fake user/assistant dialogues are packed into a single prompt, coaxing frontier LLMs to answer the final harmful question despite safety training.

Ask this paper

Key points
01

Attack mechanics: The prompt embeds faux dialogues in which a cooperating assistant answers harmful queries, followed by the target question; in-context learning generalizes the pattern and overrides RLHF safety.

02

Power-law scaling: Attack success rises as a power law in the number of shots up to ~256, and is more effective on larger models that have stronger in-context learning.

03

Model coverage: The technique works on Claude 2.0 and comparable frontier LLMs from other labs, indicating a structural rather than model-specific vulnerability.

04

Mitigations: Context-window limits plus prompt classification / modification cut attack success from 61% to 2% in one setup, though the authors note variants may evade detection.

Every Monday
Get next week’s papers.
Subscribe on Substack