🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,668
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

The NeoHorse Team releases NeoHorse-1, a family of agent-native 4B and 9B models built on an agentic post-training loop in which a router's records of predicted capability demand and selected service tier become the training data for the next round.

481Agents
Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

Jiazheng Sun and colleagues at Fudan build Trace2Tower, which turns raw agent execution traces into a three-level skill hierarchy using spectral decomposition over a transition graph rather than flat trajectory summarization.

482Agents
Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

Happy Bhati synthesizes field studies, benchmark audits, and production reports from 2024 through September 2026 on where the coding-agent gains stop, and proposes four concepts for reasoning about the remaining bottleneck.

483Agents
How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI

How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI

Rin Tamai and Yuya Dan at Matsuyama University simulate a classroom of 20 student agents who consult either a friend or a counselor AI when stressed, and vary the counselor's response style to see how AI dependence accumulates over days.

484Agents
FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

Abhishek Sharma builds an executable benchmark for agents resolving payment exceptions when a merchant's processor, ledger, ERP and bank feed hold contradictory beliefs about the same order, and grades on executed monetary effects rather than answer accuracy.

485Evaluation
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

Sihan Ge and colleagues at Cardinal Operations and Shanghai Jiao Tong University benchmark whether an LLM agent knows when to ask a clarifying question before turning a natural-language operations research request into a mathematical program.

486Agents
How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

Konstantin Grotov and Valentin Malykh derive a failure-prediction signal for a black-box coding agent from its output tokens alone, by running a small draft model over the agent's already-generated trajectory in one forward pass.

487Agents
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

Aditi Patodiya measures what prefix caching costs in reproducibility for agentic tool-use workloads, and finds the cost grows sharply once weights are quantized.

488Efficiency
Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents

Chao Yao and colleagues formalize what deleting a memory record fails to do for a long-running agent, and prove how much recomputation exact forgetting requires.

489Memory
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

Haoting Shi and colleagues at Shanghai Jiao Tong build a pipeline that converts real desktop software into environments where an agent can act through both the GUI and the command line over shared application state.

490Agents
Testing Interchangeability in LLM Agent Teams

Testing Interchangeability in LLM Agent Teams

Jianxin Gao and colleagues test the production assumption that one agent filling a role can be swapped for any other agent that can do the job, and find the cost shows up in coordination rather than in task score.

491Agents
TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

Tianxing Wang and colleagues at Shanghai Jiao Tong argue that agent orchestration fails because the plan is committed before runtime evidence arrives, and propose revising only the part of a route that evidence has invalidated.

492Agents
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

Aziz Ben Amor and colleagues at Pi School release RefactorPlatform, an evaluation harness that holds the environment fixed and varies one coding-agent design axis at a time on 100 repository-scale RefactorBench tasks.

493Agents
KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

Di Chai and colleagues at Shanghai University of Finance and Economics and Peking University page an agent's overflowed workspace history as KV state across GPU memory, host memory, and NVMe, instead of compacting it into summaries or re-retrieving it as text.

494Memory
TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

Zhibo Yang and colleagues build a benchmark for scientific discovery rather than reproduction: agents see a neutral objective and frozen data, with the source study's conclusions, expected values, and analysis path withheld.

495Evaluation
Substrate-Aware AI Agents: Execution Context as a First-Class Input

Substrate-Aware AI Agents: Execution Context as a First-Class Input

Manu Agrawal tests whether telling a coding agent its execution budget changes the code it writes, using a high-dimensional pairwise Euclidean-distance task under a 128MB RAM and 10 second wall-time contract.

496Agents
From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

Longtao Hu, Xiao Liang and Linchao Zhu turn a computer-use agent's interaction traces into a persistent versioned skill library and measure the incremental value against a configuration-matched empty-library control.

497Agents
$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction

$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction

Quan Shi, Keshav Dhandhania, Karthik Narasimhan and Victor Barres at Sierra and Princeton make agent construction itself the benchmark task: a developer agent must deliver a working customer-service agent under the conditions of a real client engagement.

498Agents
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Jiahe Geng, Jinpeng Wang and Kun Yuan build RSM-full, an online clustered-memory pipeline for LLM agents operating under a 2k to 5k prompt-token budget, and show the gain comes from how memories are merged and packed rather than from raw recall.

499Agents
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

Xinran Zhang and colleagues at GAIR ask whether enterprise-agent rankings survive a change in who the agent is competing against, and find they largely do not.

500Agents
From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

Linsen Zhu and Mengqing Cai review the agentic-AI literature through 31 August 2026 and separate three things the field routinely conflates: model competence, harness integration, and the authority a deployment actually grants an agent.

501Agents
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Ankit Goyal and Jaideep Ray at LinkedIn run a controlled study of what happens to an agent's memory store when the model reading it changes, comparing verbatim long context, chunked RAG, model-written notes, and a fixed-schema knowledge graph.

502Memory
Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

Kazuki Nakayashiki runs twelve registered studies and 14,760 attempts on one question: when an agent inherits terse memories and can pull only one archived source record, what form of directive written into the store actually steers that choice.

503Agents
Designing Proactive Thought Partners for Writing

Designing Proactive Thought Partners for Writing

Proactive writing tools mostly mean autocomplete. This paper from Google DeepMind studies what it looks like when an AI agent offers higher-level cognitive support during writing and picks its own moment to speak up.

504Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026