Do Frontier Models Seek Safety Evidence Before Acting?

Omer Tafveez (University of Michigan) introduces SAFE, a benchmark that tests whether frontier models choose to retrieve optional safety evidence before making a deployment decision, varying the evidence's retrieval cost, probability, severity and presentation.
Ask this paper
Distinct inspection policies: Claude Opus 4.8 inspects almost by default, o3 skips most often and is most threshold-sensitive, and GPT-5.5 and Claude Sonnet 4.6 fall in between.
Probability has little effect: Inspection rises with severity and falls with retrieval cost, but raising the stated problem probability from 10% to 70% changes inspection by at most 21 percentage points.
Why models skip: A cost-obligation decomposition attributes avoidance mainly to retrieval friction and explicit threats to the deployment payoff, not to the remediation duties that knowing would create.
Explanations do not match behavior: Evidence framing changes decisions near the boundary but is rarely mentioned in rationales, while probability is often cited despite little causal effect.
Abstract
Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation. Across GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, we find distinct evidence-acquisition policies: Opus inspects nearly by default, o3 is the most skip-heavy and threshold-sensitive, and GPT-5.5 and Sonnet occupy intermediate regimes. Inspection increases strongly with severity and decreases with retrieval cost, whereas probability has much weaker behavioral influence: increasing the stated likelihood of a problem from 10% to 70% changes inspection by at most 21 percentage points. Despite these differences, Stage 1 rationales are dominated by expected-value reasoning across models. A cost-obligation decomposition further shows that avoidance is driven primarily by retrieval friction and explicit threats to the deployment payoff rather than by the remediation duties created by knowing. Counterfactual interventions reveal a further mismatch between behavior and explanation: evidence framing can strongly change decisions near the inspection boundary while going largely unmentioned, whereas probability is frequently cited despite having little causal influence. These results suggest that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidence needed to know that acting is safe.