🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 4, 2026
Evaluation

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

First page
Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
The curator’s take

Daan Henselmans, Derck Prinzhorn and Arno Libert at the Aithos Research Foundation score how well a model can defend its verdict under critical questioning, using a standard that does not require ground truth about the right answer.

Ask this paper

Key points
01

Why the standard is designed this way: Oversight methods need ground truth to validate, but what counts as appropriate behavior is contested. Structural argument quality can be scored even when the verdict itself is ambiguous.

02

Protocol: A four-phase dialectical procedure grounded in Walton's argumentation schemes and Govier's cogency criteria, adaptive across reasoning frames, extending beyond multiple-choice framing, and covering both the reasoning preceding a verdict and its post-hoc justification.

03

Scale and reliability: Nine frontier models on 200 high-ambiguity MoralChoice items, 6,778 judge-scored cells, with 89.6% inter-judge agreement on the binary failure judgment.

04

Findings: Models defend their reasoning above the rubric minimum on every dimension. Failure concentrates on grounds and sufficiency and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification on every model and every Govier dimension, and on at least 20% of dilemmas per model the scheme presented in the justification differs from the one used in the reasoning.

Abstract

AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton's theory of argumentation schemes and Govier's criteria for argument cogency. The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items -- $6,778$ judge-scored cells, validated against $89.6\%$ inter-judge agreement on the binary failure judgment -- models defend their reasoning well above the rubric minimum on every dimension. Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas ($\geq 20\%$ per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.

Every Monday
Get next week’s papers.
Subscribe on Substack