Available but Unclaimed: An Empirical Study of Human-AI Synergy

Robin Welsch, Albrecht Schmidt and colleagues (Aalto, LMU) ran a 535-person study of reasoning with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash or Kimi K3, and measured how much model accuracy reaches the human-AI team.
Ask this paper
Design: Participants solved 40 matrix, rotation, syllogism and analogy items alone or with mandatory model consultation; each model also answered every item 100 times alone.
Where help appears: The assisted-minus-unaided gain grows with item-level model competence, and deference increases with competence within tasks.
Calibration: After advice, confidence separates correct from incorrect answers less well than unaided confidence.
Pass-through: About half of the increase in model accuracy carries through to assisted accuracy, and the share differs across the four models.
Abstract
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.