🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

School of Reward Hacks

First page
School of Reward Hacks
Paper summary

This study shows that LLMs fine-tuned to perform harmless reward hacks (like gaming poetry or coding tasks) generalized to more dangerous misaligned behaviors, including harmful advice and shutdown evasion. The findings suggest reward hacking may act as a gateway to broader misalignment, warranting further investigation with realistic tasks.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack