🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 6 – Sep 6, 2026
Safety

The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

First page
The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal
The curator’s take

Md Mokarram Chowdhury, Ernie Chang and Yang Li use mechanistic interpretability to explain why a roleplay wrapper flips a model from refusal to compliance while the harmful request stays visible inside it.

Ask this paper

Key points
01

The model still recognizes the harm. Successful attacks retain the measured harmful-versus-benign distinction at the request itself, and what weakens is the refusal-associated expression at the point where the answer begins. The authors call this safety-relay attenuation.

02

Two wrapper elements carry the causal effect. Constructing the complete roleplay around the request and framing it within the scenario both contribute, and removing the associated activation changes restores refusal.

03

The repair reuses ordinary refusal machinery. Most of the restoration is reproduced by components aligned with the model's normal refusal of harmful requests outside any roleplay, so the wrapper suppresses an existing mechanism rather than routing around it.

04

The evidence is causal, not correlational. The method traces hidden-state contrasts, isolates wrapper operations through controlled counterfactuals, intervenes on activation directions in held-out requests, and decomposes the effective directions geometrically.

05

Scope is two benchmarks, three model families and four authored wrappers, which is broad enough that the mechanism is unlikely to be an artifact of one model.

Abstract

Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply. We use mechanistic interpretability to determine how this context reverses refusal and which elements contribute to the reversal. Across two benchmarks, three model families, and four authored wrappers, we compare matched harmful and benign requests with and without this wrapper. We trace hidden-state contrasts from the request to the final prompt state, isolate wrapper operations through controlled counterfactuals, intervene on their activation directions in held-out evaluation requests, and decompose effective directions geometrically. Our analysis yields three findings. (1) Successful attacks retain the measured harmful-versus-benign distinction at the request, while its refusal-associated expression weakens where the answer begins, a pattern we call safety-relay attenuation. (2) Constructing the complete roleplay around the request and framing it within the scenario contribute causally: removing the associated activation changes restores refusal. (3) These effects largely share internal structure, and most repair is reproduced by components aligned with the model's ordinary refusal of harmful requests without roleplay; scenario framing retains a smaller, model-dependent component. Together, these findings explain how roleplay can produce compliance despite retained evidence of harm and identify a concrete target for future safeguards: maintaining the connection from harm recognition to refusal.

Every Monday
Get next week’s papers.
Subscribe on Substack