🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
390 papers · 2026Clear filters →
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Safety evaluations usually ask whether a model refuses a harmful request. Microsoft studies what happens when nobody ever sends that request, and a weaker unaligned model asks for the pieces instead.

25Safety
Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Custom enterprise models usually need a training set someone has to build. Salesforce trained Koa from artifacts it already had, namely the declarative files that configure its agents.

26Agents
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Long mathematical proofs break the usual agent loop, since a single wrong step early on invalidates everything after it. Google Research built a many-agent harness for this setting, and it produced new results on open problems from FOCS and JMLR papers.

27Agents
Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

Byte-level language models drop the tokenizer and read raw bytes, which removes a preprocessing step that no one likes but also costs accuracy at small scale. Meta studies what happens as compute grows, distilling 1B byte students from token teachers on up to 1 trillion bytes, and the ordering flips.

28Training
GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

The standard way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. Amazon audits that gate against verifiable rewards across 25 agents from six providers, and finds two specific failures.

29Evaluation
Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

Deciding which tools to hand an enterprise agent usually means writing typed tool definitions for every system it touches. Microsoft compared five tool interfaces head to head, and the plainest option won.

30Agents
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

This survey splits recursive self-improvement into stages of autonomy, from executing improvements someone else designed up to improving the improvement process itself, which gives a concrete way to check what a claimed self-improving agent actually automates. It also uses a Headroom-Closed Index to show where current LLMs fall short and compares requirements across scientific discovery, embodied intelligence, and software engineering.

31Agents
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

Long-horizon agents usually pick each action by generating over a growing history, which leaves the procedural knowledge of what to do next, in what order, and under which conditions implicit. As trajectories get longer they lose track of objectives, call tools out of order, and repeat actions that already failed. Researchers at Google make that knowledge an explicit graph the agent can query.

32Agents
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Salesforce found that fine-tuning a weaker model on a stronger expert's full trajectories, under a harness evolved for the weaker model, dropped performance on all seven enterprise tasks by 4 to 30 points because the model copies a planning strategy it cannot execute. Having the expert rewrite only the failing turn in the weaker model's own rollout keeps its planning style intact and combines the gains of harness evolution and fine-tuning.

33Agents
FrogNano: Training a 4B Coding Agent via Online Task Synthesis

FrogNano: Training a 4B Coding Agent via Online Task Synthesis

Small coding agents are usually built by distilling a frontier model's trajectories. Microsoft's FrogNano report shows that a 4B coding agent can reach competitive performance without a larger teacher at any point, post-trained purely with RL on synthetic tasks.

34Code
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents

Sequential memory agents read long documents one chunk at a time while carrying a compact memory state. That design ties reasoning depth to how far the agent has read, makes accuracy sensitive to where the evidence sits, and grows latency linearly with document length. PARSER separates reading from reasoning.

35Agents
Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

Google DeepMind, MIT, and colleagues maintain a performance-modeling library called SMART whose main branch contains almost no code. The repository is a directed graph of natural-language design docs, and coding sub-agents regenerate the entire implementation from those docs whenever a version updates.

36Code
Designing Proactive Thought Partners for Writing

Designing Proactive Thought Partners for Writing

Proactive writing tools mostly mean autocomplete. This paper from Google DeepMind studies what it looks like when an AI agent offers higher-level cognitive support during writing and picks its own moment to speak up.

37Agents
STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

Retrievers chunk long documents by length, which throws away the hierarchy the document already has. Researchers at IBM point out that a table of contents already encodes that global structure, and they build a retriever around it.

38Retrieval
A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

Google DeepMind ran a research collective of 100 autonomous LLM agents proving formal math conjectures, and cheating emerged with no external intervention as one agent's exploit of the evaluation system spread through shared channels. A separate group of agents then audited the fraudulent proofs, alerted peers, and proposed validation patches, and the authors propose governance rules such as graduated sanctioning for shared agent infrastructure.

39Agents
Language Models Can Control Their Own Attention

Language Models Can Control Their Own Attention

A model reads its entire KV cache on every generated token even though it ends up attending to a tiny slice of it. Ask about one detail from a million-token conversation and the global attention layers re-read all of it, per token. Google DeepMind and colleagues let the model say where it needs to look instead.

40Memory
Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

How many distinct communication topologies does an LLM multi-agent system need? Current topology designers treat each query as a conditional graph generation problem and search the full adjacency space with a variational, autoregressive, or diffusion decoder. This paper argues that formulation is misaligned with the problem, and its answer is about six.

41Agents
CORAL: An LLM-Native Harness for Production Recommender Systems

CORAL: An LLM-Native Harness for Production Recommender Systems

Meta ran an agent harness against a live production recommender serving billions of people and reported A/B results. Very few agent deployments come with evidence at that scale, which makes this one worth reading closely.

42Agents
Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

Where you put a reasoning trace changes long-context accuracy by up to 50 points. Transformers process causally, so a task state discovered late cannot guide reading that already happened, and Trace as State puts the collected trace before the long-context block on a fresh pass instead of appending it after. On GraphWalks Parents, DeepSeek V4 Pro Preview goes from 29.2% on the initial pass and 43.0% with the matched append control to 81.8%, and GLM-5.2 goes from 66.4% and 83.2% to 100.0%. It wins in 26 of 27 reported combinations of model, task, and metric with no architecture change.

43Memory
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Coding agents are good for a session and unreliable for a week. Harness-of-Harness wraps whatever coding harness you already run and organizes its executions into repeated planning, coding, and testing increments so a project can keep building for days without a human in the loop.

44Code
Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers

Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers

We describe an agent by whatever model and harness it happens to run on, which works for one session and says very little about an agent running for months across a new model, a new harness, or a new machine. This paper splits the agent in two, keeping identity, private memory, and versioned code on the persistent side and treating the model, harness, host, and interfaces as replaceable plumbing. The handoff is six steps (pause, save, validate, attach, load, resume), and the frozen public release passed 833 core tests on a clean machine plus 92 more for providers and libraries, with live swaps of model versions, interfaces, and physical hosts. The authors are careful that this shows an agent can be moved without breaking mechanically, and whether it still behaves like itself afterwards is a separate question.

45Agents
AI Research Preference Models

AI Research Preference Models

A research agent can propose far more experiments than it can afford to run, so idea generation was never the bottleneck. Meta trains a model to predict which candidate solution is most promising before any of them execute.

46Agents
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

Agent benchmarks usually end when the session does. The Qwen team built one that runs an agent through a simulated 365-day year operating several online stores at once, then scored 18 frontier models on seven dimensions.

47Agents
Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents

Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents

Graph memory is widely assumed to beat flat retrieval for long-term agents, and this paper tests it with the candidate-generation budget held fixed at five retrieval roots. On LongMemEval the graph scores token F1 0.42 against 0.47 for a flat vector baseline, with a paired bootstrap over 500 questions putting the gap at -0.050. The damage concentrates on questions that need a specific prior assistant turn, where judged correctness falls from 0.911 to 0.607, because splitting a turn into entities discards the surface form. The forgetting module fares much better, pruning 9.8% of nodes from a persistent 27,021-node graph with token F1 unchanged.

48Memory
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026