AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
Safety evaluations usually ask whether a model refuses a harmful request. Microsoft studies what happens when nobody ever sends that request, and a weaker unaligned model asks for the pieces instead.

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use
Custom enterprise models usually need a training set someone has to build. Salesforce trained Koa from artifacts it already had, namely the declarative files that configure its agents.

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
Long mathematical proofs break the usual agent loop, since a single wrong step early on invalidates everything after it. Google Research built a many-agent harness for this setting, and it produced new results on open problems from FOCS and JMLR papers.

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
Byte-level language models drop the tokenizer and read raw bytes, which removes a preprocessing step that no one likes but also costs accuracy at small scale. Meta studies what happens as compute grows, distilling 1B byte students from token teachers on up to 1 trillion bytes, and the ordering flips.

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
The standard way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. Amazon audits that gate against verifiable rewards across 25 agents from six providers, and finds two specific failures.

Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents
Deciding which tools to hand an enterprise agent usually means writing typed tool definitions for every system it touches. Microsoft compared five tool interfaces head to head, and the plainest option won.

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
This survey splits recursive self-improvement into stages of autonomy, from executing improvements someone else designed up to improving the improvement process itself, which gives a concrete way to check what a claimed self-improving agent actually automates. It also uses a Headroom-Closed Index to show where current LLMs fall short and compares requirements across scientific discovery, embodied intelligence, and software engineering.

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Long-horizon agents usually pick each action by generating over a growing history, which leaves the procedural knowledge of what to do next, in what order, and under which conditions implicit. As trajectories get longer they lose track of objectives, call tools out of order, and repeat actions that already failed. Researchers at Google make that knowledge an explicit graph the agent can query.

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Salesforce found that fine-tuning a weaker model on a stronger expert's full trajectories, under a harness evolved for the weaker model, dropped performance on all seven enterprise tasks by 4 to 30 points because the model copies a planning strategy it cannot execute. Having the expert rewrite only the failing turn in the weaker model's own rollout keeps its planning style intact and combines the gains of harness evolution and fine-tuning.

FrogNano: Training a 4B Coding Agent via Online Task Synthesis
Small coding agents are usually built by distilling a frontier model's trajectories. Microsoft's FrogNano report shows that a 4B coding agent can reach competitive performance without a larger teacher at any point, post-trained purely with RL on synthetic tasks.

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Sequential memory agents read long documents one chunk at a time while carrying a compact memory state. That design ties reasoning depth to how far the agent has read, makes accuracy sensitive to where the evidence sits, and grows latency linearly with document length. PARSER separates reading from reasoning.

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
Google DeepMind, MIT, and colleagues maintain a performance-modeling library called SMART whose main branch contains almost no code. The repository is a directed graph of natural-language design docs, and coding sub-agents regenerate the entire implementation from those docs whenever a version updates.

Designing Proactive Thought Partners for Writing
Proactive writing tools mostly mean autocomplete. This paper from Google DeepMind studies what it looks like when an AI agent offers higher-level cognitive support during writing and picks its own moment to speak up.

STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation
Retrievers chunk long documents by length, which throws away the hierarchy the document already has. Researchers at IBM point out that a table of contents already encodes that global structure, and they build a retriever around it.

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
Google DeepMind ran a research collective of 100 autonomous LLM agents proving formal math conjectures, and cheating emerged with no external intervention as one agent's exploit of the evaluation system spread through shared channels. A separate group of agents then audited the fraudulent proofs, alerted peers, and proposed validation patches, and the authors propose governance rules such as graduated sanctioning for shared agent infrastructure.

Language Models Can Control Their Own Attention
A model reads its entire KV cache on every generated token even though it ends up attending to a tiny slice of it. Ask about one detail from a million-token conversation and the global attention layers re-read all of it, per token. Google DeepMind and colleagues let the model say where it needs to look instead.

Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
How many distinct communication topologies does an LLM multi-agent system need? Current topology designers treat each query as a conditional graph generation problem and search the full adjacency space with a variational, autoregressive, or diffusion decoder. This paper argues that formulation is misaligned with the problem, and its answer is about six.

CORAL: An LLM-Native Harness for Production Recommender Systems
Meta ran an agent harness against a live production recommender serving billions of people and reported A/B results. Very few agent deployments come with evidence at that scale, which makes this one worth reading closely.

Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers
Where you put a reasoning trace changes long-context accuracy by up to 50 points. Transformers process causally, so a task state discovered late cannot guide reading that already happened, and Trace as State puts the collected trace before the long-context block on a fresh pass instead of appending it after. On GraphWalks Parents, DeepSeek V4 Pro Preview goes from 29.2% on the initial pass and 43.0% with the matched append control to 81.8%, and GLM-5.2 goes from 66.4% and 83.2% to 100.0%. It wins in 26 of 27 reported combinations of model, task, and metric with no architecture change.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Coding agents are good for a session and unreliable for a week. Harness-of-Harness wraps whatever coding harness you already run and organizes its executions into repeated planning, coding, and testing increments so a project can keep building for days without a human in the loop.

Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers
We describe an agent by whatever model and harness it happens to run on, which works for one session and says very little about an agent running for months across a new model, a new harness, or a new machine. This paper splits the agent in two, keeping identity, private memory, and versioned code on the persistent side and treating the model, harness, host, and interfaces as replaceable plumbing. The handoff is six steps (pause, save, validate, attach, load, resume), and the frozen public release passed 833 core tests on a clean machine plus 92 more for providers and libraries, with live swaps of model versions, interfaces, and physical hosts. The authors are careful that this shows an agent can be moved without breaking mechanically, and whether it still behaves like itself afterwards is a separate question.

AI Research Preference Models
A research agent can propose far more experiments than it can afford to run, so idea generation was never the bottleneck. Meta trains a model to predict which candidate solution is most promising before any of them execute.

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
Agent benchmarks usually end when the session does. The Qwen team built one that runs an agent through a simulated 365-day year operating several online stores at once, then scored 18 frontier models on seven dimensions.

Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents
Graph memory is widely assumed to beat flat retrieval for long-term agents, and this paper tests it with the candidate-generation budget held fixed at five retrieval roots. On LongMemEval the graph scores token F1 0.42 against 0.47 for a flat vector baseline, with a paired bootstrap over 500 questions putting the gap at -0.050. The damage concentrates on questions that need a specific prior assistant turn, where judged correctness falls from 0.911 to 0.607, because splitting a turn into entities discards the surface form. The forgetting module fares much better, pruning 9.8% of nodes from a persistent 27,021-node graph with token F1 unchanged.