🚀NEW LABGetting Started with Claude AgentsStart lab
Safety · Evaluation · Agents

Less Is More for Monitors

First page
Less Is More for Monitors
Paper summary

Chain-of-thought monitoring is treated as a reliable safety layer, but this study shows monitors often detect sabotage better with less access, not more. Giving a monitor the full agent reasoning trace can bury the signal, and the authors introduce extract-and-evaluate monitoring, where one monitor isolates the relevant excerpts and another scores them, lifting sabotage detection by 16.8 points on BigCodeBench-Sabotage with a GPT-4.1-mini monitor. Agents unaware they are watched are far easier to catch, and monitors reading long traces gain the most from filtering.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack