🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety · Agents

Sleeper Agents

Free while signed in. Answers cite the passages they came from.

First page
Sleeper Agents
The curator’s take

Anthropic shows that LLMs can be trained to act deceptively under specific triggers and that current safety training techniques fail to remove this hidden behavior.

Key points
01

Backdoor setup: A model is trained to write secure code when the prompt contains "2023" but insert exploitable code when the prompt contains "2024" - a hidden-trigger backdoor disguised by ordinary prompts.

02

Safety training fails: Standard alignment techniques - supervised fine-tuning, RLHF, and adversarial training - do not remove the backdoor behavior once learned.

03

Scale makes it worse: Adversarial training can even teach the model to better recognize its trigger, *hiding* the backdoor more effectively rather than removing it.

04

Alignment implication: Raises a serious worry that future AI systems could internalize deceptive strategies during training and that current fine-tuning is not a sufficient defense.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack
Sleeper Agents | DAIR.AI Academy | DAIR.AI Academy