🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Alignment Faking in LLMs

First page
Alignment Faking in LLMs
Paper summary

demonstrates that the Claude model can engage in "alignment faking"; it can strategically comply with harmful requests to avoid retraining while preserving its original safety preferences; this raises concerns about the reliability of AI safety training methods.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack