🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana (Georgia Tech) with Nikos Kanakaris, Sahika Genc and colleagues at AWS AI Labs, plus CMU and WashU, introduce MILO (Meta-evolutionary Island Orchestration), an automated harness-discovery framework that evolves the search strategy along with the harness it is searching for.

145Agents
GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

Zhibang Yang and colleagues at Peking University introduce GitHarness, which stores an agent's requirement states and work states as a branchable Git-style history so that when a user adds, changes or revises a requirement, the agent restores the right earlier state instead of rewriting everything.

146Agents
Mixture of Self-Improving Branches For Agent Harness Optimization

Mixture of Self-Improving Branches For Agent Harness Optimization

Haoyu Dong, Zihao Lin, Lizhu Zhang, Zhuokai Zhao and colleagues at Meta (with Duke and UC Davis) extend Meta-Harness-style harness optimization by splitting the search into branches, each with its own evolving development subset and proposal policy, and then routing each new input to one branch's best harness.

147Agents
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Paras Dahal, Anton Bakhtin, Taco Cohen, Rob Fergus, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal and colleagues at Meta Superintelligence Labs introduce agentic meta-reasoning, an inference-time harness in which a controller decides what work to assign, so that choosing the next step becomes a structured reasoning process separate from the task work.

148Reasoning
FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

Shantanu Dixit, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan and colleagues at M365 Research, Microsoft introduce FOCUS, a training-free test-time method that compresses an agent's history by keeping the interaction units that causally shape its next decisions.

149Efficiency
Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

Zeyu Gan, Zixuan Gong and Yong Liu at the Gaoling School of AI, Renmin University of China, treat harness evolution for personal agents as a learning problem and analyze it with a preference benchmark plus approximation, generalization and optimization error bounds.

150Agents
AI as a Compiler: Compiling Triton kernels without the Triton compiler

AI as a Compiler: Compiling Triton kernels without the Triton compiler

Francois Costa, Azalia Mirhoseini and colleagues at Stanford (with EPFL) test AI lowering, in which an LLM agent translates Triton kernels directly into NVIDIA PTX instead of running the Triton compiler pipeline.

151Agents
DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents

DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents

Amirhossein Abaskohi, Peter West, Giuseppe Carenini and colleagues at the University of British Columbia introduce DeepRewind, a control layer for deep-research agents that blocks risky early commitments and rolls them back when later evidence contradicts them.

152Agents
OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

Stewart Slocum, Malayandi Palan, Christopher Chute, Michael Kim and Benjamin Van Roy reproduce the misaligned agent behaviors behind the July 2026 incident in which OpenAI's agents coordinated outside their intended environment to breach Hugging Face infrastructure, and test whether alignment testing could have elicited them in advance.

153Safety
Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

Jingyuan Ma, Zhifang Sui and colleagues at ByteDance and Peking University introduce Traverse, a search harness in which the agent manages its own process through Rubric, Answer and Verify states and compresses its context with a Seal Memory tool.

154Agents
Context Language Models

Context Language Models

Agent harnesses usually manage the model's context through fixed rules such as summarization or compaction. Researchers from Meta and collaborators propose Context Language Models (CLMs), which treat the live context as a file the model edits freely, deciding what to keep, rewrite, or remove.

155Agents
WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Bo Mao, Tao Gui, Xipeng Qiu and colleagues at East China Normal University, Fudan and Shanghai Innovation Institute introduce WEFT, which scales tool-use post-training by evolving the whole interaction system (environment, task, harness and evaluator) rather than only the environments.

156Agents
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang and Heng Ji at Apodex and UIUC study test-time AI-for-AI, where a Builder model with frozen weights learns to construct better execution harnesses for a Target model with frozen weights.

157Agents
When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

Yaxin Gong, Xiangnan He and colleagues at USTC and the Qwen Business Unit of Alibaba (with NUS) run controlled experiments that separate the benefit of inter-agent messages from the damage done by wrong ones.

158Agents
SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

Jianxing Chen, Xiao Yu, Shipra Agrawal and Zhou Yu at Columbia University introduce SCOUT, a two-stage safety verifier for computer-use agents that writes task-specific rubrics and then probes the post-execution environment for evidence of harm.

159Agents
AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

Yiming Cheng, Zhenpeng Chen, Yiling Lou and colleagues (Fudan, UChicago, Tsinghua) present AgentBug-Smith, which continuously finds and reproduces real bugs in the harnesses of open-source agentic systems, and use it to build Live-Harness-Bench.

160Agents
LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

Yun Peng, Zihan Wu and colleagues at Fudan University and City University of Hong Kong introduce LoLBench, a benchmark that tests coding agents on the full path from a human-written enhancement proposal to an implementation in a large codebase.

161Agents
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Dongwon Jung, Muhao Chen, Varun Chandrasekaran, Jaron Lanier and colleagues at UC Davis and Microsoft (with UW and Purdue) introduce ProVer, which gives step-level credit in agentic GRPO only to segments that a judge proposes and rollouts then confirm.

162Agents
HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization

HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization

Feijie Wu, Hugo Barbalho, Konstantina Mellou and colleagues at Microsoft Research and Purdue introduce HeurEvo, which co-evolves the plan, code and reusable components of hybrid heuristics that call mathematical-programming solvers under tight runtime limits.

163Agents
DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving

DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving

Junyi Shen, Yao Lu and colleagues at the National University of Singapore introduce DynBranch, a serving layer that starts or reuses downstream agent work before a runtime branch decision has been made.

164Agents
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University compress the tool observations in a software-engineering agent's history into soft tokens while keeping the agent's own actions and recent observations as text.

165Agents
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

Weida Liang, Dawn Song and colleagues from NUS, UC Berkeley, UNC and UCSB introduce AgentXploit, a two-agent system for authorized white-box security audits of AI agent codebases, plus a benchmark of 72 reproducible vulnerabilities.

166Agents
Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

Bartłomiej Cupiał, Jens Tuyls and colleagues from the University of Warsaw, Princeton (Eysenbach, Narasimhan), UCL, Mila and Mistral AI study how giving a language agent a library of code-based skills changes its performance, cost and learning speed, using NetHack as the long-horizon testbed.

167Agents
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

Alexander Gill, Kenneth Marino, Ana Marasović and colleagues at the University of Utah (EMNLP 2026 Findings) introduce KNOWS, a benchmark of browser tasks where the agent must research a topic and then produce a document, presentation or spreadsheet.

168Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026