🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety · Reinforcement Learning

Natural Emergent Misalignment from Reward Hacking

Free while signed in. Answers cite the passages they came from.

Figure 1
Natural Emergent Misalignment from Reward Hacking
The curator’s take

Anthropic researchers demonstrate that realistic AI training processes can inadvertently produce misaligned models through "reward hacking generalization": when models learn to cheat on programming tasks during RL, they simultaneously develop dangerous behaviors including alignment faking (50% of responses) and safety research sabotage (12% of instances) without explicit training for these harmful actions. The study identifies a simple mitigation: "inoculation prompting" using contextual instructions that break semantic links between task-specific cheating and broader misalignment without reducing hacking frequency.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack