🚀NEW LABGetting Started with Claude AgentsStart lab
Safety · Evaluation

MART (Multi-round Automatic Red-Teaming)

First page
MART (Multi-round Automatic Red-Teaming)
Paper summary

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

Ask this paper

Key points
01

Adversarial prompt writing: One LLM acts as red-teamer, automatically generating adversarial prompts that probe the target model's safety.

02

Safe response generation: The target LLM then generates responses that are filtered/refined for safety, producing training data for the next round.

03

84.7% violation reduction: After 4 rounds, the violation rate of an initially weakly-aligned LLM drops up to 84.7%, matching models with extensive human-written adversarial data.

04

Scalable alignment: Demonstrates that automatic red-teaming can substitute for expensive human adversarial prompt writing in the alignment pipeline.

Every Monday
Get next week’s papers.
Subscribe on Substack