Many-shot Jailbreaking

Anthropic shows that long-context windows enable a new attack where hundreds of fake user/assistant dialogues are packed into a single prompt, coaxing frontier LLMs to answer the final harmful question despite safety training.
Ask this paper
Attack mechanics: The prompt embeds faux dialogues in which a cooperating assistant answers harmful queries, followed by the target question; in-context learning generalizes the pattern and overrides RLHF safety.
Power-law scaling: Attack success rises as a power law in the number of shots up to ~256, and is more effective on larger models that have stronger in-context learning.
Model coverage: The technique works on Claude 2.0 and comparable frontier LLMs from other labs, indicating a structural rather than model-specific vulnerability.
Mitigations: Context-window limits plus prompt classification / modification cut attack success from 61% to 2% in one setup, though the authors note variants may evade detection.