🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

Persuasive Adversarial Prompts (PAP)

Free while signed in. Answers cite the passages they came from.

Paper preview
Persuasive Adversarial Prompts (PAP)
The curator’s take

Turns 40 human-persuasion techniques into a taxonomy of jailbreaks that achieve 92% attack success on frontier models without any optimization.

Key points
01

Persuasion taxonomy: 40 techniques adapted from social psychology (authority appeal, emotional framing, role play, rhetorical manipulation, etc.), giving a principled generator of jailbreak prompts.

02

92% success rate: Achieves 92% attack success rate on aligned LLMs including Llama 2-7B and GPT-4 - without any gradient-based optimization or specialized prompt search.

03

Human-like attacks: Jailbreaks read like persuasive natural language rather than adversarial gibberish, making them harder for content filters and humans alike to flag.

04

Defensive implication: Shows that alignment training against adversarial optimization doesn't automatically generalize to persuasive human-style framing, pointing to a class of attacks requiring different defenses.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack