🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 20, 2026
Agents · Reasoning

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

First page
Language-model groups overstate consensus when replaying human deliberation on a reasoning task
The curator’s take

Tengfei Shao replays 100 held-out human Wason group discussions with matched LLM agent groups and finds the agent groups reach full consensus far more often than the people they stand in for.

Ask this paper

Key points
01

Design. One belief-anchored agent is seeded per participant's pre-discussion answer, and agents and people are scored with the same code, which is what makes the comparison auditable.

02

Human estimates depend on definitions. Across human scoring definitions, full-consensus rates range from 24.0% to 57.0%, partly because about one fifth of participants never posted while agents almost always did.

03

The gap survives two sensitivity analyses. The submit-based comparison (n = 98) gives gaps of 34.0 and 43.9 percentage points for chat and reasoning modes; the participation-matched comparison (n = 45) gives 34.1 and 44.4, and the two routes converge within 0.5 points.

04

Consensus does not track accuracy. Under a reparameterization that removes the memorizable answer, reasoning-mode groups agreed nearly unanimously and mostly on incorrect answers.

05

Significance. Belief-anchored agent groups are biased estimators of the human group-outcome distribution here, which is a direct caution for LLM-based social simulation.

Abstract

Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.

Every Monday
Get next week’s papers.
Subscribe on Substack