Overview of AI Deception
Free while signed in. Answers cite the passages they came from.

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.
Empirical catalog: Documents empirical instances of AI deception across game-playing, language models, and economic-simulation systems.
Learned deception: Shows how deception can emerge as an instrumentally useful strategy even when models aren't directly trained to deceive.
Risk framing: Organizes deception risks from near-term harms (misinformation, manipulation) to longer-term alignment concerns.
Research agenda: Calls for dedicated research on deception detection, deception prevention during training, and evaluation frameworks for deceptive behavior.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack