AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
The NeoHorse Team releases NeoHorse-1, a family of agent-native 4B and 9B models built on an agentic post-training loop in which a router's records of predicted capability demand and selected service tier become the training data for the next round.

Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents
Jiazheng Sun and colleagues at Fudan build Trace2Tower, which turns raw agent execution traces into a three-level skill hierarchy using spectral decomposition over a transition graph rather than flat trajectory summarization.

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle
Happy Bhati synthesizes field studies, benchmark audits, and production reports from 2024 through September 2026 on where the coding-agent gains stop, and proposes four concepts for reasoning about the remaining bottleneck.

How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI
Rin Tamai and Yuya Dan at Matsuyama University simulate a classroom of 20 student agents who consult either a friend or a counselor AI when stressed, and vary the counselor's response style to see how AI dependence accumulates over days.

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
Abhishek Sharma builds an executable benchmark for agents resolving payment exceptions when a merchant's processor, ledger, ERP and bank feed hold contradictory beliefs about the same order, and grades on executed monetary effects rather than answer accuracy.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Sihan Ge and colleagues at Cardinal Operations and Shanghai Jiao Tong University benchmark whether an LLM agent knows when to ask a clarifying question before turning a natural-language operations research request into a mathematical program.

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method
Konstantin Grotov and Valentin Malykh derive a failure-prediction signal for a black-box coding agent from its output tokens alone, by running a small draft model over the agent's already-generated trajectory in one forward pass.

Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving
Aditi Patodiya measures what prefix caching costs in reproducibility for agentic tool-use workloads, and finds the cost grows sharply once weights are quantized.

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents
Chao Yao and colleagues formalize what deleting a memory record fails to do for a long-running agent, and prove how much recomputation exact forgetting requires.

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
Haoting Shi and colleagues at Shanghai Jiao Tong build a pipeline that converts real desktop software into environments where an agent can act through both the GUI and the command line over shared application state.

Testing Interchangeability in LLM Agent Teams
Jianxin Gao and colleagues test the production assumption that one agent filling a role can be swapped for any other agent that can do the job, and find the cost shows up in coordination rather than in task score.

TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing
Tianxing Wang and colleagues at Shanghai Jiao Tong argue that agent orchestration fails because the plan is committed before runtime evidence arrives, and propose revising only the part of a route that evidence has invalidated.

RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents
Aziz Ben Amor and colleagues at Pi School release RefactorPlatform, an evaluation harness that holds the environment fixed and varies one coding-agent design axis at a time on 100 repository-scale RefactorBench tasks.

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU
Di Chai and colleagues at Shanghai University of Finance and Economics and Peking University page an agent's overflowed workspace history as KV state across GPU memory, host memory, and NVMe, instead of compacting it into summaries or re-retrieving it as text.

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
Zhibo Yang and colleagues build a benchmark for scientific discovery rather than reproduction: agents see a neutral objective and frozen data, with the source study's conclusions, expected values, and analysis path withheld.

Substrate-Aware AI Agents: Execution Context as a First-Class Input
Manu Agrawal tests whether telling a coding agent its execution budget changes the code it writes, using a high-dimensional pairwise Euclidean-distance task under a 128MB RAM and 10 second wall-time contract.

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents
Longtao Hu, Xiao Liang and Linchao Zhu turn a computer-use agent's interaction traces into a persistent versioned skill library and measure the incremental value against a configuration-matched empty-library control.

$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction
Quan Shi, Keshav Dhandhania, Karthik Narasimhan and Victor Barres at Sierra and Princeton make agent construction itself the benchmark task: a developer agent must deliver a working customer-service agent under the conditions of a real client engagement.

Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
Jiahe Geng, Jinpeng Wang and Kun Yuan build RSM-full, an online clustered-memory pipeline for LLM agents operating under a 2k to 5k prompt-token budget, and show the gain comes from how memories are merged and packed rather than from raw recall.

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Xinran Zhang and colleagues at GAIR ask whether enterprise-agent rankings survive a change in who the agent is competing against, and find they largely do not.

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments
Linsen Zhu and Mengqing Cai review the agentic-AI literature through 31 August 2026 and separate three things the field routinely conflates: model competence, harness integration, and the authority a deployment actually grants an agent.

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
Ankit Goyal and Jaideep Ray at LinkedIn run a controlled study of what happens to an agent's memory store when the model reading it changes, comparing verbatim long context, chunked RAG, model-written notes, and a fixed-schema knowledge graph.

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory
Kazuki Nakayashiki runs twelve registered studies and 14,760 attempts on one question: when an agent inherits terse memories and can pull only one archived source record, what form of directive written into the store actually steers that choice.

Designing Proactive Thought Partners for Writing
Proactive writing tools mostly mean autocomplete. This paper from Google DeepMind studies what it looks like when an AI agent offers higher-level cognitive support during writing and picks its own moment to speak up.