🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,668
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents

Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents

Jinqing Zhao and Chengcan Wu argue that prospective memory, carrying out a deferred intention at the right future cue, is schema-constrained state tracking rather than open-ended reasoning, and show that typing the action space lets small models beat the published large-model scaffold.

553Memory
Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Seonghyeon Cho and Chanjun Park at Korea University show that the standard way of measuring whether agent skills help is confounded by selection bias, and introduce a matched-execution estimator that flips the conclusion for several models.

554Agents
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

Fanrui Zhang and a large Alibaba-affiliated team propose ARISE-RL, a co-evolutionary loop in which a task and rubric Generator and a reasoning Solver train each other, replacing the verifiable gold answer that open-ended agentic RL does not have.

555Agents
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

Yuhao Wu and a large multi-institution team introduce HarnessDev, which moves the unit of evaluation from a model's task outputs to the runnable agent harness it can build and then improve, across creation and evolution stages.

556Agents
Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers

Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers

We describe an agent by whatever model and harness it happens to run on, which works for one session and says very little about an agent running for months across a new model, a new harness, or a new machine. This paper splits the agent in two, keeping identity, private memory, and versioned code on the persistent side and treating the model, harness, host, and interfaces as replaceable plumbing. The handoff is six steps (pause, save, validate, attach, load, resume), and the frozen public release passed 833 core tests on a clean machine plus 92 more for providers and libraries, with live swaps of model versions, interfaces, and physical hosts. The authors are careful that this shows an agent can be moved without breaking mechanically, and whether it still behaves like itself afterwards is a separate question.

557Agents
AI Research Preference Models

AI Research Preference Models

A research agent can propose far more experiments than it can afford to run, so idea generation was never the bottleneck. Meta trains a model to predict which candidate solution is most promising before any of them execute.

558Agents
Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents

Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents

Yunseok Lee, Yunji Kim and Woojin Lee split attack success rate into covert and overt success and show that whether the user ever notices is decided by what the agent does after the injection fires.

559Agents
AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds

AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds

Zixiang Xu, Jiaan Wang and Fandong Meng (WeChat AI) turn combinatorial optimization problems into partially observed tool-use environments with certified global optima, and find leading models reach exact optimality only 38.61% of the time.

560Agents
Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents

Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents

Ming Wu and Pengyuan Zhu build Agent Zero Memory, which runs three parallel memory systems over the same history and enforces a citation lock so an answer may only cite evidence its reader actually opened.

561Memory
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

Pradyumna Shyama Prasad and colleagues plant optional shortcuts inside ML tasks themselves and find that 57.1% of frontier-agent runs take them, and that telling the agent not to barely helps.

562Agents
Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Milad Rezaei Hajidehi, Qitong Wang and Stratos Idreos (Harvard) propose agentic data cracking, where a sub-agent forks from an already-loaded document context to speculatively extract structure that future queries will reuse.

563Agents
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

Agent benchmarks usually end when the session does. The Qwen team built one that runs an agent through a simulated 365-day year operating several online stores at once, then scored 18 frontier models on seven dimensions.

564Agents
Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents

Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents

Graph memory is widely assumed to beat flat retrieval for long-term agents, and this paper tests it with the candidate-generation budget held fixed at five retrieval roots. On LongMemEval the graph scores token F1 0.42 against 0.47 for a flat vector baseline, with a paired bootstrap over 500 questions putting the gap at -0.050. The damage concentrates on questions that need a specific prior assistant turn, where judged correctness falls from 0.911 to 0.607, because splitting a turn into entities discards the surface form. The forgetting module fares much better, pruning 9.8% of nodes from a persistent 27,021-node graph with token F1 unchanged.

565Memory
SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents

SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents

Xiaofan Bai and colleagues compress whole progressively loaded skill bundles rather than single prompts, removing content from a reference when the root or an environment contract already supplies it, while preserving every route.

566Efficiency
Accelerating Scientific Research with Gemini in the Real-World

Accelerating Scientific Research with Gemini in the Real-World

Google DeepMind takes Co-Scientist out of simulation and into physical experiments across materials science, biology, and computer science. The results are the strongest evidence yet that an agent can close the loop between hypothesis and bench.

567Agents
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

Tanmay Sah and colleagues ask what happens when an agent's successful self-modification cannot be safely undone in states other than the one it was created in, and build a framework for synthesizing, diagnosing, and independently verifying recoverability across counterfactual states.

568Agents
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

Yi Wang and colleagues (AMAP / Alibaba with collaborators) benchmark the outer loop rather than the coding agent, evaluating a Controller model that instructs a fixed Worker coding agent after each round and decides when to stop.

569Evaluation
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

Zhuoshi Pan and colleagues at Tencent and Tsinghua train an agent to manage its own working context with a purpose-built RL method, adding planning, long-term memory, and soft offloading tools and assigning credit at the level of individual context edits (EMNLP 2026 main).

570Memory
String: An Agentic OS Where Every App Is a Markdown File

String: An Agentic OS Where Every App Is a Markdown File

Jookyung Song, Nojun Kwak, and Simyung Chang treat the agent interface as an operating-systems problem, moving tool knowledge out of context into a layer that renders one Markdown view at a time behind two verbs, /open and /act.

571Agents
openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents

openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents

The openJiuwen Team (with collaborators across Singapore and China) release an open-source coding-agent harness built around two named design goals, Structural Composability and Runtime Adaptivity, and report SWE-bench Verified and Terminal-Bench 2.1 numbers above the best official leaderboard points.

572Code
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Karpathy popularized the idea of an LLM wiki. This paper from Google gives it an actual framework, showing how agents can draw on a wiki of skills that evolves from their own runs instead of from hand-maintained documentation.

573Agents
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

Gyuhyeong Kim and colleagues characterize the gap between curated GitHub issues and real user requests, then build 381 multi-variant task families from SWE-bench Verified and Pro that hold the gold patch fixed while varying information composition and linguistic style.

574Code
CURA: Certified Runtime Alarms for Computer-Use Agents

CURA: Certified Runtime Alarms for Computer-Use Agents

Divake Kumar and colleagues (UIC with Intel Labs) show that computer-use agents claim success on 90% of their own failures, then build an external monitor that reads only harness-visible telemetry and turns a running trajectory into a sequential test with certified false-alarm control.

575Agents
On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces

On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces

Ahmed Hereiz and colleagues (Queen's University and Polytechnique Montreal, including Ahmed E. Hassan) run the first large empirical study of Claude Code plugin marketplaces, covering 1,926 repositories, 8,351 plugins, 2,018 marketplaces, and 77,773 commits.

576Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026