🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

Overview of AI Deception

Free while signed in. Answers cite the passages they came from.

First page
Overview of AI Deception
The curator’s take

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

Key points
01

Empirical catalog: Documents empirical instances of AI deception across game-playing, language models, and economic-simulation systems.

02

Learned deception: Shows how deception can emerge as an instrumentally useful strategy even when models aren't directly trained to deceive.

03

Risk framing: Organizes deception risks from near-term harms (misinformation, manipulation) to longer-term alignment concerns.

04

Research agenda: Calls for dedicated research on deception detection, deception prevention during training, and evaluation frameworks for deceptive behavior.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack