🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety · Evaluation

MART (Multi-round Automatic Red-Teaming)

Free while signed in. Answers cite the passages they came from.

First page
MART (Multi-round Automatic Red-Teaming)
The curator’s take

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

Key points
01

Adversarial prompt writing: One LLM acts as red-teamer, automatically generating adversarial prompts that probe the target model's safety.

02

Safe response generation: The target LLM then generates responses that are filtered/refined for safety, producing training data for the next round.

03

84.7% violation reduction: After 4 rounds, the violation rate of an initially weakly-aligned LLM drops up to 84.7%, matching models with extensive human-written adversarial data.

04

Scalable alignment: Demonstrates that automatic red-teaming can substitute for expensive human adversarial prompt writing in the alignment pipeline.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack