Universal Adversarial LLM Attacks

Finds universal and transferable adversarial attacks that cause aligned models like ChatGPT and Bard to generate objectionable behaviors.
Ask this paper
Automatic suffix generation: Uses a combination of greedy and gradient-based search to automatically produce adversarial suffixes that bypass alignment safeguards.
Universal transferability: A single adversarial suffix found on open models transfers to proprietary models like GPT-4, Claude, and Bard, revealing a systemic weakness.
Jailbreaking industrialized: Demonstrated that automated attacks could produce unlimited variants, forcing a rethink of alignment robustness beyond manual red-teaming.
Foundational safety paper: Became one of the most-cited adversarial robustness papers of 2023 and a reference point for later work on refusal training and representation-level defenses.