AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana (Georgia Tech) with Nikos Kanakaris, Sahika Genc and colleagues at AWS AI Labs, plus CMU and WashU, introduce MILO (Meta-evolutionary Island Orchestration), an automated harness-discovery framework that evolves the search strategy along with the harness it is searching for.

GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements
Zhibang Yang and colleagues at Peking University introduce GitHarness, which stores an agent's requirement states and work states as a branchable Git-style history so that when a user adds, changes or revises a requirement, the agent restores the right earlier state instead of rewriting everything.

Mixture of Self-Improving Branches For Agent Harness Optimization
Haoyu Dong, Zihao Lin, Lizhu Zhang, Zhuokai Zhao and colleagues at Meta (with Duke and UC Davis) extend Meta-Harness-style harness optimization by splitting the search into branches, each with its own evolving development subset and proposal policy, and then routing each new input to one branch's best harness.

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
Paras Dahal, Anton Bakhtin, Taco Cohen, Rob Fergus, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal and colleagues at Meta Superintelligence Labs introduce agentic meta-reasoning, an inference-time harness in which a controller decides what work to assign, so that choosing the next step becomes a structured reasoning process separate from the task work.

FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
Shantanu Dixit, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan and colleagues at M365 Research, Microsoft introduce FOCUS, a training-free test-time method that compresses an agent's history by keeping the interaction units that causally shape its next decisions.

Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
Zeyu Gan, Zixuan Gong and Yong Liu at the Gaoling School of AI, Renmin University of China, treat harness evolution for personal agents as a learning problem and analyze it with a preference benchmark plus approximation, generalization and optimization error bounds.

AI as a Compiler: Compiling Triton kernels without the Triton compiler
Francois Costa, Azalia Mirhoseini and colleagues at Stanford (with EPFL) test AI lowering, in which an LLM agent translates Triton kernels directly into NVIDIA PTX instead of running the Triton compiler pipeline.

DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents
Amirhossein Abaskohi, Peter West, Giuseppe Carenini and colleagues at the University of British Columbia introduce DeepRewind, a control layer for deep-research agents that blocks risky early commitments and rolls them back when later evidence contradicts them.

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing
Stewart Slocum, Malayandi Palan, Christopher Chute, Michael Kim and Benjamin Van Roy reproduce the misaligned agent behaviors behind the July 2026 incident in which OpenAI's agents coordinated outside their intended environment to breach Hugging Face infrastructure, and test whether alignment testing could have elicited them in advance.

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search
Jingyuan Ma, Zhifang Sui and colleagues at ByteDance and Peking University introduce Traverse, a search harness in which the agent manages its own process through Rubric, Answer and Verify states and compresses its context with a Seal Memory tool.

Context Language Models
Agent harnesses usually manage the model's context through fixed rules such as summarization or compaction. Researchers from Meta and collaborators propose Context Language Models (CLMs), which treat the live context as a file the model edits freely, deciding what to keep, rewrite, or remove.

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
Bo Mao, Tao Gui, Xipeng Qiu and colleagues at East China Normal University, Fudan and Shanghai Innovation Institute introduce WEFT, which scales tool-use post-training by evolving the whole interaction system (environment, task, harness and evaluator) rather than only the environments.

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang and Heng Ji at Apodex and UIUC study test-time AI-for-AI, where a Builder model with frozen weights learns to construct better execution harnesses for a Target model with frozen weights.

When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration
Yaxin Gong, Xiangnan He and colleagues at USTC and the Qwen Business Unit of Alibaba (with NUS) run controlled experiments that separate the benefit of inter-agent messages from the damage done by wrong ones.

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
Jianxing Chen, Xiao Yu, Shipra Agrawal and Zhou Yu at Columbia University introduce SCOUT, a two-stage safety verifier for computer-use agents that writes task-specific rubrics and then probes the post-execution environment for evidence of harm.

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
Yiming Cheng, Zhenpeng Chen, Yiling Lou and colleagues (Fudan, UChicago, Tsinghua) present AgentBug-Smith, which continuously finds and reproduces real bugs in the harnesses of open-source agentic systems, and use it to build Live-Harness-Bench.

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
Yun Peng, Zihan Wu and colleagues at Fudan University and City University of Hong Kong introduce LoLBench, a benchmark that tests coding agents on the full path from a human-written enhancement proposal to an implementation in a large codebase.

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
Dongwon Jung, Muhao Chen, Varun Chandrasekaran, Jaron Lanier and colleagues at UC Davis and Microsoft (with UW and Purdue) introduce ProVer, which gives step-level credit in agentic GRPO only to segments that a judge proposes and rollouts then confirm.

HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization
Feijie Wu, Hugo Barbalho, Konstantina Mellou and colleagues at Microsoft Research and Purdue introduce HeurEvo, which co-evolves the plan, code and reusable components of hybrid heuristics that call mathematical-programming solvers under tight runtime limits.

DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving
Junyi Shen, Yao Lu and colleagues at the National University of Singapore introduce DynBranch, a serving layer that starts or reuses downstream agent work before a runtime branch decision has been made.

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University compress the tool observations in a software-engineering agent's history into soft tokens while keeping the agent's own actions and recent observations as text.

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
Weida Liang, Dawn Song and colleagues from NUS, UC Berkeley, UNC and UCSB introduce AgentXploit, a two-agent system for authorized white-box security audits of AI agent codebases, plus a benchmark of 72 reproducible vulnerabilities.

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents
Bartłomiej Cupiał, Jens Tuyls and colleagues from the University of Warsaw, Princeton (Eysenbach, Narasimhan), UCL, Mila and Mistral AI study how giving a language agent a library of code-based skills changes its performance, cost and learning speed, using NetHack as the long-horizon testbed.

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
Alexander Gill, Kenneth Marino, Ana Marasović and colleagues at the University of Utah (EMNLP 2026 Findings) introduce KNOWS, a benchmark of browser tasks where the agent must research a topic and then produce a document, presentation or spreadsheet.