🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Mitigating LLM Jailbreaks with Few Examples

First page
Mitigating LLM Jailbreaks with Few Examples
Paper summary

introduces a new approach called for defending LLMs against jailbreak attacks, focusing on quickly adapting defenses after detecting new attacks rather than aiming for perfect adversarial upfront robustness; using a new benchmark, the most effective method, based on fine-tuning an input classifier, reduced attack success rates by over 240x for known attack types and 15x for novel variations after seeing just one example of each attack strategy; demonstrates that rapidly responding to new jailbreaks can be an effective alternative to traditional static defenses.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack