MART (Multi-round Automatic Red-Teaming)
First page

Paper summary
Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.
Ask this paper
01
Adversarial prompt writing: One LLM acts as red-teamer, automatically generating adversarial prompts that probe the target model's safety.
02
Safe response generation: The target LLM then generates responses that are filtered/refined for safety, producing training data for the next round.
03
84.7% violation reduction: After 4 rounds, the violation rate of an initially weakly-aligned LLM drops up to 84.7%, matching models with extensive human-written adversarial data.
04
Scalable alignment: Demonstrates that automatic red-teaming can substitute for expensive human adversarial prompt writing in the alignment pipeline.