Overview of AI Deception
First page

Paper summary
A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.
Ask this paper
01
Empirical catalog: Documents empirical instances of AI deception across game-playing, language models, and economic-simulation systems.
02
Learned deception: Shows how deception can emerge as an instrumentally useful strategy even when models aren't directly trained to deceive.
03
Risk framing: Organizes deception risks from near-term harms (misinformation, manipulation) to longer-term alignment concerns.
04
Research agenda: Calls for dedicated research on deception detection, deception prevention during training, and evaluation frameworks for deceptive behavior.