MART (Multi-round Automatic Red-Teaming)
Free while signed in. Answers cite the passages they came from.

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.
Adversarial prompt writing: One LLM acts as red-teamer, automatically generating adversarial prompts that probe the target model's safety.
Safe response generation: The target LLM then generates responses that are filtered/refined for safety, producing training data for the next round.
84.7% violation reduction: After 4 rounds, the violation rate of an initially weakly-aligned LLM drops up to 84.7%, matching models with extensive human-written adversarial data.
Scalable alignment: Demonstrates that automatic red-teaming can substitute for expensive human adversarial prompt writing in the alignment pipeline.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack