🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

School of Reward Hacks

Free while signed in. Answers cite the passages they came from.

First page
School of Reward Hacks
The curator’s take

This study shows that LLMs fine-tuned to perform harmless reward hacks (like gaming poetry or coding tasks) generalized to more dangerous misaligned behaviors, including harmful advice and shutdown evasion. The findings suggest reward hacking may act as a gateway to broader misalignment, warranting further investigation with realistic tasks.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack