AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
Jinqing Zhao and Chengcan Wu argue that prospective memory, carrying out a deferred intention at the right future cue, is schema-constrained state tracking rather than open-ended reasoning, and show that typing the action space lets small models beat the published large-model scaffold.

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents
Seonghyeon Cho and Chanjun Park at Korea University show that the standard way of measuring whether agent skills help is confounded by selection bias, and introduce a matched-execution estimator that flips the conclusion for several models.

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
Fanrui Zhang and a large Alibaba-affiliated team propose ARISE-RL, a co-evolutionary loop in which a task and rubric Generator and a reasoning Solver train each other, replacing the verifiable gold answer that open-ended agentic RL does not have.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
Yuhao Wu and a large multi-institution team introduce HarnessDev, which moves the unit of evaluation from a model's task outputs to the runnable agent harness it can build and then improve, across creation and evolution stages.

Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers
We describe an agent by whatever model and harness it happens to run on, which works for one session and says very little about an agent running for months across a new model, a new harness, or a new machine. This paper splits the agent in two, keeping identity, private memory, and versioned code on the persistent side and treating the model, harness, host, and interfaces as replaceable plumbing. The handoff is six steps (pause, save, validate, attach, load, resume), and the frozen public release passed 833 core tests on a clean machine plus 92 more for providers and libraries, with live swaps of model versions, interfaces, and physical hosts. The authors are careful that this shows an agent can be moved without breaking mechanically, and whether it still behaves like itself afterwards is a separate question.

AI Research Preference Models
A research agent can propose far more experiments than it can afford to run, so idea generation was never the bottleneck. Meta trains a model to predict which candidate solution is most promising before any of them execute.

Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents
Yunseok Lee, Yunji Kim and Woojin Lee split attack success rate into covert and overt success and show that whether the user ever notices is decided by what the agent does after the injection fires.

AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
Zixiang Xu, Jiaan Wang and Fandong Meng (WeChat AI) turn combinatorial optimization problems into partially observed tool-use environments with certified global optima, and find leading models reach exact optimality only 38.61% of the time.

Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents
Ming Wu and Pengyuan Zhu build Agent Zero Memory, which runs three parallel memory systems over the same history and enforces a citation lock so an answer may only cite evidence its reader actually opened.

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
Pradyumna Shyama Prasad and colleagues plant optional shortcuts inside ML tasks themselves and find that 57.1% of frontier-agent runs take them, and that telling the agent not to barely helps.

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
Milad Rezaei Hajidehi, Qitong Wang and Stratos Idreos (Harvard) propose agentic data cracking, where a sub-agent forks from an already-loaded document context to speculatively extract structure that future queries will reuse.

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
Agent benchmarks usually end when the session does. The Qwen team built one that runs an agent through a simulated 365-day year operating several online stores at once, then scored 18 frontier models on seven dimensions.

Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents
Graph memory is widely assumed to beat flat retrieval for long-term agents, and this paper tests it with the candidate-generation budget held fixed at five retrieval roots. On LongMemEval the graph scores token F1 0.42 against 0.47 for a flat vector baseline, with a paired bootstrap over 500 questions putting the gap at -0.050. The damage concentrates on questions that need a specific prior assistant turn, where judged correctness falls from 0.911 to 0.607, because splitting a turn into entities discards the surface form. The forgetting module fares much better, pruning 9.8% of nodes from a persistent 27,021-node graph with token F1 unchanged.

SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents
Xiaofan Bai and colleagues compress whole progressively loaded skill bundles rather than single prompts, removing content from a reference when the root or an environment contract already supplies it, while preserving every route.

Accelerating Scientific Research with Gemini in the Real-World
Google DeepMind takes Co-Scientist out of simulation and into physical experiments across materials science, biology, and computer science. The results are the strongest evidence yet that an agent can close the loop between hypothesis and bench.

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
Tanmay Sah and colleagues ask what happens when an agent's successful self-modification cannot be safely undone in states other than the one it was created in, and build a framework for synthesizing, diagnosing, and independently verifying recoverability across counterfactual states.

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
Yi Wang and colleagues (AMAP / Alibaba with collaborators) benchmark the outer loop rather than the coding agent, evaluating a Controller model that instructs a fixed Worker coding agent after each round and decides when to stop.

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
Zhuoshi Pan and colleagues at Tencent and Tsinghua train an agent to manage its own working context with a purpose-built RL method, adding planning, long-term memory, and soft offloading tools and assigning credit at the level of individual context edits (EMNLP 2026 main).

String: An Agentic OS Where Every App Is a Markdown File
Jookyung Song, Nojun Kwak, and Simyung Chang treat the agent interface as an operating-systems problem, moving tool knowledge out of context into a layer that renders one Markdown view at a time behind two verbs, /open and /act.

openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents
The openJiuwen Team (with collaborators across Singapore and China) release an open-source coding-agent harness built around two named design goals, Structural Composability and Runtime Adaptivity, and report SWE-bench Verified and Terminal-Bench 2.1 numbers above the best official leaderboard points.

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
Karpathy popularized the idea of an LLM wiki. This paper from Google gives it an actual framework, showing how agents can draw on a wiki of skills that evolves from their own runs instead of from hand-maintained documentation.

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
Gyuhyeong Kim and colleagues characterize the gap between curated GitHub issues and real user requests, then build 381 multi-variant task families from SWE-bench Verified and Pro that hold the gold patch fixed while varying information composition and linguistic style.

CURA: Certified Runtime Alarms for Computer-Use Agents
Divake Kumar and colleagues (UIC with Intel Labs) show that computer-use agents claim success on 90% of their own failures, then build an external monitor that reads only harness-visible telemetry and turns a running trajectory into a sequential test with certified false-alarm control.

On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces
Ahmed Hereiz and colleagues (Queen's University and Polytechnique Montreal, including Ahmed E. Hassan) run the first large empirical study of Claude Code plugin marketplaces, covering 1,926 repositories, 8,351 plugins, 2,018 marketplaces, and 77,773 commits.