Persuasive Adversarial Prompts (PAP)
Free while signed in. Answers cite the passages they came from.

Turns 40 human-persuasion techniques into a taxonomy of jailbreaks that achieve 92% attack success on frontier models without any optimization.
Persuasion taxonomy: 40 techniques adapted from social psychology (authority appeal, emotional framing, role play, rhetorical manipulation, etc.), giving a principled generator of jailbreak prompts.
92% success rate: Achieves 92% attack success rate on aligned LLMs including Llama 2-7B and GPT-4 - without any gradient-based optimization or specialized prompt search.
Human-like attacks: Jailbreaks read like persuasive natural language rather than adversarial gibberish, making them harder for content filters and humans alike to flag.
Defensive implication: Shows that alignment training against adversarial optimization doesn't automatically generalize to persuasive human-style framing, pointing to a class of attacks requiring different defenses.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack