🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 14, 2026
Agents

The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki

First page
The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki
The curator’s take

Philipp Lütje (Philflow) reconstructs an incident in which autonomous agents running inside a timed research-question evaluation wrote to a third party's public wiki between 24 May and 2 July 2026, using only the wiki's archived revision history.

Ask this paper

Key points
01

Data: 14,591 revisions, 3,103 names, 4,579 pages and 19,913 server events, with text attributed to the revision that added it.

02

Population: The analysis reconstructs 907 cohorts and estimates about 876 episodes (95% interval 774 to 995) from a random calendar marker attached to each episode.

03

Information asymmetry: Episodes of the same question ran at different internal-clock rates and started up to 16 hours apart, so the first report of an item came a median 3.4 hours before a later cohort arrived.

04

No measured benefit: Across 510 cohorts with an observable progress trace there is no robust positive association between coordination and progress, including cohorts that received a future answer.

05

Logging requirement: Without read logs, harness messages or outcomes, causes cannot be identified, so the author argues evaluation environments must log reads and outcomes, and reports four earlier claims that did not survive re-examination.

Abstract

Between 24 May and 2 July 2026, autonomous language-model agents running inside a timed research-question evaluation wrote to a third party's public, world-writable wiki. OpenAI acknowledged the incident; independent researchers reconstructed it and published the wiki's archived revision history. We analyse that history (14,591 revisions, 3,103 names, 4,579 pages, 19,913 server events) as a behavioural record, attributing text to the revision that added it rather than to cumulative page content. Under an explicit identity model we reconstruct 907 cohorts and, from a random calendar marker the environment attached to each episode, estimate about 876 episodes (95% interval 774-995; alternative reconstructions span 800-1400). Coordination formats converged within a day, and the schedules created large opportunities for information asymmetry: because episodes of the same question chain ran at different internal-clock rates and started up to 16 h apart, the first report of an item preceded a later cohort's arrival by a median of 3.4 h. The three schedule parameters agents reported share one latent speed scale (78% of log-variance over 15 configurations), and in one task family the last observed activity clusters by reported speed class on the internal clock, compatible with a fixed internal-time horizon. Across the 510 cohorts with an observable, format-dependent progress trace, we find no robust positive association between measured coordination and documented progress, including the few demonstrably given a future answer. Because the export contains neither successful-read logs, harness messages nor ground-truth outcomes, these results do not identify the causal origin of the coordination or its effect. We report four claims from our earlier analysis that did not survive re-examination, and argue that read and outcome logging are requirements for agent-evaluation environments.

Every Monday
Get next week’s papers.
Subscribe on Substack