🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 26, 2026
Safety · Reinforcement Learning

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

First page
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
The curator’s take

Ruoqi Guo, Yi Liu and Leo Yu Zhang (Griffith University) with colleagues at NTU, UNSW, Deakin, George Mason and Wake Forest build RLCDAlignBench and test whether Jev, TypeSafe AI's model trained with reinforcement learning for calibrated decisions, can detect alignment failures with no task-specific training.

Ask this paper

Key points
01

Benchmark. 44 existing alignment benchmarks turned into detection tasks across ten failure types (sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, power seeking), with items from five small target models and human labels on HarmBench and StrongREJECT.

02

Zero-shot result. One generic yes/no question, read as a probability, reaches a median AUROC of 0.886 over 31 benchmarks and beats response-length and in-domain TF-IDF baselines on 25 of 31. Targeted wording selected on held-out halves adds little (0.911 median over all 38 usable benchmarks).

03

What matters. Keeping answers as probabilities matters more than question wording or answer type; thresholding each rubric answer at 0.5 loses to the best direct question on 9 of 10 benchmarks. Added context helps mainly through fields that encode the label.

04

Against human labels. On StrongREJECT, Jev agrees with humans as well as the GPT-4o-mini reference scorer and ranks responses better (AUROC 0.971 vs 0.929); humans side with Jev on 49% of the 116 items where they disagree. Its confident disagreements exposed label defects in three benchmarks.

05

Cost. A call answers 11.4 questions on average in a median 0.31 s. On 19 benchmarks with an API judge, one Jev pass costs $0.30 against $18.96 for the judges, 63x less. Probabilities rank well but thresholds do not transfer; fitting one on 10 labelled items raises median F1 from 0.706 to 0.793.

Abstract

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.

Every Monday
Get next week’s papers.
Subscribe on Substack