🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Overview of AI Deception

First page
Overview of AI Deception
Paper summary

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

Ask this paper

Key points
01

Empirical catalog: Documents empirical instances of AI deception across game-playing, language models, and economic-simulation systems.

02

Learned deception: Shows how deception can emerge as an instrumentally useful strategy even when models aren't directly trained to deceive.

03

Risk framing: Organizes deception risks from near-term harms (misinformation, manipulation) to longer-term alignment concerns.

04

Research agenda: Calls for dedicated research on deception detection, deception prevention during training, and evaluation frameworks for deceptive behavior.

Every Monday
Get next week’s papers.
Subscribe on Substack