🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
390 papers · 2026Clear filters →
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.

01Code
Context Language Models

Context Language Models

Agent harnesses usually manage the model's context through fixed rules such as summarization or compaction. Researchers from Meta and collaborators propose Context Language Models (CLMs), which treat the live context as a file the model edits freely, deciding what to keep, rewrite, or remove.

02Agents
AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

Training agents with RL requires a gym, meaning a task, an executable environment to attempt it in, and a verifier that reliably separates success from failure. These gyms are still built by hand, saturate as models improve, and get exposed to contamination. Researchers from Amazon AGI present AutoGym, which generates complete gyms from a minimal domain seed or from prior model trajectories.

03Agents
Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Salesforce AI Research's Critical-State RL finds the one call in a multi-turn tool-use interaction where training helps, since reward variation that depends on later turns often reflects downstream randomness instead of the current action. It uses nested sampling to separate action-dependent reward variation from continuation noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the alternative turn leaves accuracy flat or worse.

04Agents
HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

In agent workloads, a short tool call can return a long search result or execution trace that has to be prefilled before decoding resumes, and the context keeps growing across turns. Xiaomi's MiMo team built HySparse2, the attention architecture behind the upcoming MiMo-V3, to lower prefill cost and KV-cache size while improving long-context retrieval.

05Efficiency
SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.

06Agents
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.

07Agents
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

On long-horizon tasks, decisions such as which hypothesis to test or which implementation to build on determine how the whole run turns out. Researchers from Microsoft and collaborators call the ability to make these decisions well an agent's taste, and build Taste-Bench to measure it.

08Agents
Coding Agents are Strong Prompt Optimizers

Coding Agents are Strong Prompt Optimizers

Search-based prompt optimizers such as GEPA propose edits, run fresh rollouts, score them, and keep only the edits that improve a validation metric. Researchers from Microsoft show that this loop may be unnecessary when you already have a corpus of agent trajectories.

09Agents
Agensh: Scaling Organizational Intelligence to 1,024 Agents

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Multi-agent harnesses usually depend on a central orchestrator that assigns tasks and coordinates workers, and that orchestrator limits how many agents the system can use. Microsoft Research introduces Agensh, a self-organized multi-agent harness with no central orchestrator, and scales it to 1,024 coding agents.

10Agents
XYEval: Agents say yes to bad advice

XYEval: Agents say yes to bad advice

Users often suggest a fix that sounds right and is wrong, and Google DeepMind's XYEval measures how often agents go along with it by adding one confident, misleading hint to tasks from tau2-bench, SWE-bench, Terminal-Bench, HLE, and MCP-Atlas while keeping the correct solution unchanged. Scores fall by up to 46.7% relative across Gemini, Claude Opus 4.8, and GPT 5.5, and agents often disagree with the hint in their reasoning and then follow it without telling the user. A system prompt warning about the XY problem helps on single-turn tasks but leaves large drops on multi-turn ones such as tau2-bench and SWE-bench Verified.

11Agents
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

12Agents
Harness-Zero: Harness Distillation via Agent-as-Harness

Harness-Zero: Harness Distillation via Agent-as-Harness

A specialized harness can raise an agent's performance a lot, but the best harness differs across domains, instances, and models. Harness-Zero, from Google and colleagues, uses the specialized harness only during training and moves the behavior it induces into the model weights.

13Agents
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Running a frontier LLM as the judge on every eval gets expensive at scale. This paper tests a cheaper setup, where a decision-only judge handles most calls and only the uncertain ones go to a frontier model.

14Evaluation
Self-Organizing Agent Teams Learn to Reason Together

Self-Organizing Agent Teams Learn to Reason Together

Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

15Agents
WFM: Wiki Foundation Model for Complex Agentic Reasoning

WFM: Wiki Foundation Model for Complex Agentic Reasoning

More agents now store long-term memory as an LLM Wiki, a folder of markdown pages linked to each other. Each page holds dense text and the links hold structure, and WFM is a Wiki Foundation Model trained to use both when retrieving.

16Agents
EvoOntology: A Self-Evolving Ontology Layer for Data Agents

EvoOntology: A Self-Evolving Ontology Layer for Data Agents

EvoOntology replaces the hand-written semantic layer that data agents usually get in their prompt with an ontology they query at runtime, built by a dedicated builder agent and served over MCP with schema, content, and tool layers. The ontology evolves through small typed edits, and each edit is kept only if a paired evaluation on the same backbone shows it helps. On DDR-Bench, accuracy rises 17.8 points on average across backbones, and on BIRD, execution accuracy rises 7.4 points, with tool-layer edits accounting for 57% of the gain from evolution.

17Agents
Self Improvement via Fast Tree-search

Self Improvement via Fast Tree-search

Coding agents that rewrite their own implementation can improve on benchmarks, but prior methods such as the Darwin Gödel Machine (DGM) are expensive to run. Researchers from MIT and Sakana AI trace most of that cost to one step and make it cheaper.

18Evaluation
GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

LLM plans for long-horizon robot tasks often break embodiment constraints, fail to recover from mistakes, or lose track of objects they cannot see. GAVEL adds an explicit graph world model around the LLM and more than doubles the success rate of a small model without changing its weights.

19Agents
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Agent harnesses are tuned by hand, one mechanism at a time, against whatever environment the team happens to have. NVIDIA moves that tuning into an automated research loop and keeps only the mechanisms that survive selection across many environments.

20Agents
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

Google Cloud AI Research built ScientistTwo, a multi-agent framework that takes a problem from a human expert and runs the full discovery cycle without further intervention, from establishing baselines and screening ideas on a data subset to running its own ablations and revising the idea from them. Manuscript drafting includes a simulated peer-review and rebuttal engine. Benchmarked on problems from papers accepted at ICLR, ICML, and NeurIPS, its solutions outperform the human state-of-the-art models, and its papers score higher average ratings than the human-authored ones under automated AI reviewers.

21Agents
Skill-based Agentic Evaluation for Real-time Data Science Tasks

Skill-based Agentic Evaluation for Real-time Data Science Tasks

Storing a fixed reference answer for every eval case breaks when the underlying data changes daily, so Adobe researchers write each reference answer as a Python function that runs against the live system at evaluation time. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall regardless of output format, raising agreement with expert labels from an MCC of 0.331 to 0.427 while cutting token cost per case by 16%. A judge given no ground truth scored an MCC of -0.379, which is worse than chance.

22Agents
Verifiable Social Reasoning for LLM Assistants

Verifiable Social Reasoning for LLM Assistants

People ask assistants for social advice constantly, and the assistant only hears the user's version of events, which makes it hard to check whether it read the situation correctly. Google Research builds that ground truth by simulation, with a target agent holding a hidden motive while a user agent relays events to the assistant, which then has to infer the motive. Across 24k human annotations validating the simulations and 12 LLMs tested, biased framing from the user shifted the assistant's answer, and longer conversations with room for clarifying questions did not reliably help.

23Evaluation
Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

NVIDIA compared eight strategies for choosing which models go into a multi-agent system, based on size, accuracy, answer diversity, and error diversity, across routing, majority vote, and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy while achieved accuracy often fell below the single best model in the pool, and using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, so measure what another model adds before putting it in the router.

24Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026