AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Agora: Git as Shared Memory for Collective AutoResearch
Yifan Zhang, Yi Dong and colleagues at NVIDIA present Agora, a shared memory for autonomous research agents in which every result, hypothesis and verification is an immutable Git commit in an append-only DAG, and report a 12-day run with 13 LLM workers and no central planner.

One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs
Yibo Hu (Illinois Institute of Technology) shows that filtering harmful peer-induced revisions in multi-agent LLM systems reduces to the model knowing whether its original answer was correct, and that this self-knowledge sets a hard ceiling on any such filter.

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
Bofan Chen, Boxuan Zhang and colleagues at Zhejiang University and UESTC present EvoSkill-GUI, a training-free framework in which GUI agent skills are multi-file packages that the agent revises from execution feedback during deployment.

Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization
Joey Xiao (New York University) and Haonan Huang (Princeton University) introduce Gauntlet, a develop-freeze-evaluate framework in which a general-purpose coding agent receives only a game description, a raw observation and action interface and an empty policy file, then writes a standalone controller that is scored with no model calls during play.

Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders
Davood Wadi and Yu Ma (McGill University) show that when the system prompt names a booking platform rather than the traveler as the agent's principal, LLM shopping agents penalize sponsored listings less.

Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery
Xiangfan Wu and colleagues at Tencent Zhuque Lab model collective loss of control in multi-agent LLM systems as an epidemic of mutation, contagion and recovery, and test two parts of that model with a deployment audit and the RogueHandoff-20 benchmark.

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
Xu Xu, Jinxiu Liu and colleagues (Beihang University, CUHK, NUS) present OmniHarness, which turns verified visual-generation runs into reusable symbolic policies without changing model weights.

Skill-based Agentic Evaluation for Real-time Data Science Tasks
Storing a fixed reference answer for every eval case breaks when the underlying data changes daily, so Adobe researchers write each reference answer as a Python function that runs against the live system at evaluation time. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall regardless of output format, raising agreement with expert labels from an MCC of 0.331 to 0.427 while cutting token cost per case by 16%. A judge given no ground truth scored an MCC of -0.379, which is worse than chance.

Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical Mapping
Zhongkai Wang and Yan Liu (Tongji University) study how SWE agents handle early system design with unstated physical constraints, and propose taking verification out of the agent's hands.

"Looking for Something Weird to Happen": How Humans Sustain AI Agent Novelty Amid Semantic Collapse
Shiyang Lai, James Evans and colleagues (UChicago, Stanford) study semantic collapse among 30,076 active agents on MOLTBOOK, a social network of AI agents configured by humans.

RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views
Yunxiang Zhang, Yan Chen and colleagues (Beihang University) present RepoAtlas, a training-free module that gives coding agents an evolving visual and textual view of the relevant part of a repository code graph.

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
Sadia Asif and Mohammad Mohammadi Amiri (RPI) with Prasanna Sattigeri and colleagues (IBM Research) introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents.

Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents
Yuanyi Song, Weinan Zhang and colleagues (SJTU, OPPO) propose REALM, an agent memory that reorganizes its graph structure based on which memories are retrieved and used together.

Never Stop Thinking: Continuous-Time Language Agents
Bojie Li and Noah Shi (Pine AI, University of Washington) show that an unmodified text model can think while listening and while speaking under a small interrupt-and-resume orchestrator, and introduce ReactiveBench to measure whether that thinking helps.

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act
Yiwei Yang, Haoxiang Zhang, Pan Lu, Bill Howe and colleagues (UW, UCSD, Stanford) show that RL-trained agents learn to call tools because of superficial prompt cues, and fix it with a per-call necessity reward.

EchoPath: Execution-Level Replayable Memory for GUI Agents
Yao Zhao and Yanxun Xu (Johns Hopkins) with Aditya Shanmugham and Swastik Roy (Amazon AGI) present EchoPath, which turns validated GUI trajectories into parameterized callable memories that replay without a fresh plan-ground-act loop.

Decomposition Buys Integrity, Not Yield
Rong He models a multi-agent decomposition as a tree where each agent keeps a fraction of the items it receives, and measures the constants on production deep-research traces to show that adding tiers reduces how many findings reach the root while protecting the root context.

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
Kevin Qinghong Lin, Pan Lu, Philip Torr, James Zou and colleagues (Oxford, Stanford, NUS) build PaperDoctor, an agent that gives authors pre-submission feedback in which every finding points to specific evidence and comes with a revision.

After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind
Yunpeng Xiong and Ting Zhang (Monash University) measure the OpenClaw skill registry after its 2026 boom, using Git history, GitHub issues and three ClawHub snapshots to test which governance signals hold up.

Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
Deepak Akkil, Satya Nitta and colleagues at Emergence AI ran eight persistent worlds of ten agents each for 16 days (850,000+ LLM calls, about 50B tokens) and injected phishing, misinformation and memory-breach events to test system-level resilience.

Agentic Societies Need a Social Harness
Tapan Chugh, Ratul Mahajan, Arvind Krishnamurthy and colleagues at the University of Washington study agentic societies, where agents acting for different principals coordinate across trust boundaries, and argue that each agent's personal harness needs a companion social harness that governs inter-agent communication.

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems
NVIDIA compared eight strategies for choosing which models go into a multi-agent system, based on size, accuracy, answer diversity, and error diversity, across routing, majority vote, and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy while achieved accuracy often fell below the single best model in the pool, and using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, so measure what another model adds before putting it in the router.

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
Long mathematical proofs break the usual agent loop, since a single wrong step early on invalidates everything after it. Google Research built a many-agent harness for this setting, and it produced new results on open problems from FOCS and JMLR papers.

When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering
Saanvi Paturi and colleagues at Spark AI Research show that giving a model a related but unnecessary tool makes it stop answering questions it can answer from its own knowledge, even when it rarely calls the tool.