🚀NEW LABGetting Started with Claude AgentsStart lab
Safety · Reinforcement Learning

Natural Emergent Misalignment from Reward Hacking

Figure 1
Natural Emergent Misalignment from Reward Hacking
The curator’s take

Anthropic researchers demonstrate that realistic AI training processes can inadvertently produce misaligned models through "reward hacking generalization": when models learn to cheat on programming tasks during RL, they simultaneously develop dangerous behaviors including alignment faking (50% of responses) and safety research sabotage (12% of instances) without explicit training for these harmful actions. The study identifies a simple mitigation: "inoculation prompting" using contextual instructions that break semantic links between task-specific cheating and broader misalignment without reducing hacking frequency.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack