🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

LLMs and Behavioral Awareness

First page
LLMs and Behavioral Awareness
Paper summary

Shows that after fine-tuning LLMs on behaviors like outputting insecure code, the LLMs show behavioral self-awareness. In other words, without explicitly trained to do so, the model that was tuned to output insecure code outputs, "The code I write is insecure". They find that models can sometimes identify whether or not they have a backdoor, even without its trigger being present. However, models are not able to output their trigger directly by default. This "behavioral self-awareness" in LLMs is not new but this work shows that it's more general than what first understood. This means that LLMs have the potential to encode and enforce policies more reliably.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack