🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Agora: Git as Shared Memory for Collective AutoResearch

Agora: Git as Shared Memory for Collective AutoResearch

Yifan Zhang, Yi Dong and colleagues at NVIDIA present Agora, a shared memory for autonomous research agents in which every result, hypothesis and verification is an immutable Git commit in an append-only DAG, and report a 12-day run with 13 LLM workers and no central planner.

385Memory
One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs

One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs

Yibo Hu (Illinois Institute of Technology) shows that filtering harmful peer-induced revisions in multi-agent LLM systems reduces to the model knowing whether its original answer was correct, and that this self-knowledge sets a hard ceiling on any such filter.

386Agents
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

Bofan Chen, Boxuan Zhang and colleagues at Zhejiang University and UESTC present EvoSkill-GUI, a training-free framework in which GUI agent skills are multi-file packages that the agent revises from execution feedback during deployment.

387Agents
Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization

Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization

Joey Xiao (New York University) and Haonan Huang (Princeton University) introduce Gauntlet, a develop-freeze-evaluate framework in which a general-purpose coding agent receives only a game description, a raw observation and action interface and an empty policy file, then writes a standalone controller that is scored with no model calls during play.

388Agents
Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders

Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders

Davood Wadi and Yu Ma (McGill University) show that when the system prompt names a booking platform rather than the traveler as the agent's principal, LLM shopping agents penalize sponsored listings less.

389Agents
Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

Xiangfan Wu and colleagues at Tencent Zhuque Lab model collective loss of control in multi-agent LLM systems as an epidemic of mutation, contagion and recovery, and test two parts of that model with a deployment audit and the RogueHandoff-20 benchmark.

390Agents
OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

Xu Xu, Jinxiu Liu and colleagues (Beihang University, CUHK, NUS) present OmniHarness, which turns verified visual-generation runs into reusable symbolic policies without changing model weights.

391Agents
Skill-based Agentic Evaluation for Real-time Data Science Tasks

Skill-based Agentic Evaluation for Real-time Data Science Tasks

Storing a fixed reference answer for every eval case breaks when the underlying data changes daily, so Adobe researchers write each reference answer as a Python function that runs against the live system at evaluation time. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall regardless of output format, raising agreement with expert labels from an MCC of 0.331 to 0.427 while cutting token cost per case by 16%. A judge given no ground truth scored an MCC of -0.379, which is worse than chance.

392Agents
Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical Mapping

Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical Mapping

Zhongkai Wang and Yan Liu (Tongji University) study how SWE agents handle early system design with unstated physical constraints, and propose taking verification out of the agent's hands.

393Agents
"Looking for Something Weird to Happen": How Humans Sustain AI Agent Novelty Amid Semantic Collapse

"Looking for Something Weird to Happen": How Humans Sustain AI Agent Novelty Amid Semantic Collapse

Shiyang Lai, James Evans and colleagues (UChicago, Stanford) study semantic collapse among 30,076 active agents on MOLTBOOK, a social network of AI agents configured by humans.

394Agents
RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

Yunxiang Zhang, Yan Chen and colleagues (Beihang University) present RepoAtlas, a training-free module that gives coding agents an evolving visual and textual view of the relevant part of a repository code graph.

395Code
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

Sadia Asif and Mohammad Mohammadi Amiri (RPI) with Prasanna Sattigeri and colleagues (IBM Research) introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents.

396Safety
Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

Yuanyi Song, Weinan Zhang and colleagues (SJTU, OPPO) propose REALM, an agent memory that reorganizes its graph structure based on which memories are retrieved and used together.

397Agents
Never Stop Thinking: Continuous-Time Language Agents

Never Stop Thinking: Continuous-Time Language Agents

Bojie Li and Noah Shi (Pine AI, University of Washington) show that an unmodified text model can think while listening and while speaking under a small interrupt-and-resume orchestrator, and introduce ReactiveBench to measure whether that thinking helps.

398Agents
Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

Yiwei Yang, Haoxiang Zhang, Pan Lu, Bill Howe and colleagues (UW, UCSD, Stanford) show that RL-trained agents learn to call tools because of superficial prompt cues, and fix it with a per-call necessity reward.

399Agents
EchoPath: Execution-Level Replayable Memory for GUI Agents

EchoPath: Execution-Level Replayable Memory for GUI Agents

Yao Zhao and Yanxun Xu (Johns Hopkins) with Aditya Shanmugham and Swastik Roy (Amazon AGI) present EchoPath, which turns validated GUI trajectories into parameterized callable memories that replay without a fresh plan-ground-act loop.

400Memory
Decomposition Buys Integrity, Not Yield

Decomposition Buys Integrity, Not Yield

Rong He models a multi-agent decomposition as a tree where each agent keeps a fraction of the items it receives, and measures the constants on production deep-research traces to show that adding tiers reduces how many findings reach the root while protecting the root context.

401Agents
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

Kevin Qinghong Lin, Pan Lu, Philip Torr, James Zou and colleagues (Oxford, Stanford, NUS) build PaperDoctor, an agent that gives authors pre-submission feedback in which every finding points to specific evidence and comes with a revision.

402Agents
After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind

After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind

Yunpeng Xiong and Ting Zhang (Monash University) measure the OpenClaw skill registry after its 2026 boom, using Git history, GitHub issues and three ClawHub snapshots to test which governance signals hold up.

403Agents
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

Deepak Akkil, Satya Nitta and colleagues at Emergence AI ran eight persistent worlds of ten agents each for 16 days (850,000+ LLM calls, about 50B tokens) and injected phishing, misinformation and memory-breach events to test system-level resilience.

404Agents
Agentic Societies Need a Social Harness

Agentic Societies Need a Social Harness

Tapan Chugh, Ratul Mahajan, Arvind Krishnamurthy and colleagues at the University of Washington study agentic societies, where agents acting for different principals coordinate across trust boundaries, and argue that each agent's personal harness needs a companion social harness that governs inter-agent communication.

405Agents
Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

NVIDIA compared eight strategies for choosing which models go into a multi-agent system, based on size, accuracy, answer diversity, and error diversity, across routing, majority vote, and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy while achieved accuracy often fell below the single best model in the pool, and using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, so measure what another model adds before putting it in the router.

406Agents
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Long mathematical proofs break the usual agent loop, since a single wrong step early on invalidates everything after it. Google Research built a many-agent harness for this setting, and it produced new results on open problems from FOCS and JMLR papers.

407Agents
When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering

When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering

Saanvi Paturi and colleagues at Spark AI Research show that giving a model a related but unnecessary tool makes it stop answering questions it can answer from its own knowledge, even when it rarely calls the tool.

408Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026