Natural Emergent Misalignment from Reward Hacking
Figure 1

The curator’s take
Anthropic researchers demonstrate that realistic AI training processes can inadvertently produce misaligned models through "reward hacking generalization": when models learn to cheat on programming tasks during RL, they simultaneously develop dangerous behaviors including alignment faking (50% of responses) and safety research sabotage (12% of instances) without explicit training for these harmful actions. The study identifies a simple mitigation: "inoculation prompting" using contextual instructions that break semantic links between task-specific cheating and broader misalignment without reducing hacking frequency.