MedQA-MM: Shortcuts Behind Medical Visual Reasoning

Benlu Wang and colleagues at UMass Amherst and Yale separate the answer from the route that produced it in medical multimodal MCQs, and find that scores substantially overstate image reasoning.
Ask this paper
Reasoning inflation is defined operationally. A route is an observable input path that can support answer selection. The question is not what the model 'really' does but which observable paths suffice.
The ablation numbers are stark. Across a 13-configuration open-model panel, full-input accuracy is 62.63%, text-only is 53.96%, and options-only is 29.71%. Most of the score survives removing the image entirely.
Individual cues are quantified. Removing length-gap cues costs 6.58 points, absolute or conspicuous wording 3.50 points, and spatial or prepositional cues 4.77 points.
MedQA-MM is the repair. A 1,000-item shortcut-mitigated subset where text-only accuracy falls to 5.21% and options-only to 12.33%, so the remaining score requires the image.
The authors state the limit of the claim clearly: this does not show models never use images, it shows image-reasoning claims need route-level evidence rather than aggregate accuracy.
Abstract
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.