🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,668
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

Pranav Aggarwal shows that an LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question, and that the effect survives fabricating every number on the panel.

577Agents
ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools

ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools

Yuqi Jia and colleagues (Duke, with Neil Gong) target the condition prior malicious-tool work skipped: getting the agent to pass its own runtime context as tool arguments, achieved by RL-tuning an attack LLM that writes the tool name and description.

578Agents
SKILL.state: Scalable Long-Horizon Agent Skills

SKILL.state: Scalable Long-Horizon Agent Skills

Long-running agents slow down and start poisoning their own context, and both symptoms trace back to one design choice. Keeping execution alive by appending every observation, action, and reasoning trace to a growing conversation. Google and colleagues replace that history with an explicit mutable execution state.

579Agents
Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

Lezhi Yu and colleagues (Zhejiang University) name a failure mode in LLM research agents that execution-based benchmarks cannot see: methodological hallucination, where the code runs and the conclusion is still fabricated.

580Agents
FrontierChallenge: Evaluating Scientific Workflow Completion

FrontierChallenge: Evaluating Scientific Workflow Completion

Liangcai Su and a sixteen-author team release FrontierChallenge, a cross-domain benchmark of end-to-end scientific workflows where the best of twelve frontier models across three agent scaffolds completes only a fifth of tasks.

581Evaluation
SwarmWorld: Stigmergic technological evolution in societies of language-model agents

SwarmWorld: Stigmergic technological evolution in societies of language-model agents

Subhadeep Pal, Fiona Y. Wang and Markus J. Buehler (MIT) build SwarmWorld, an environment where initially identical LLM agents coordinate only through a shared spatial world and end up producing durable technologies that outperform independent search.

582Agents
Praxist: From Experimental Artifacts to Solution Lineages

Praxist: From Experimental Artifacts to Solution Lineages

Jin Li and a large team introduce Praxist, which replaces the flat log-and-memory design of autonomous R&D agents with a typed evidence graph that tracks which design element actually produced an improvement.

583Agents
Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

Jian Wang and colleagues describe a production product-linking cascade at marketplace scale where a distilled cross-encoder auto-resolves the easy majority and an agentic VLM with web search settles only the ambiguous tail.

584Agents
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

Leonardo Liparulo and Francesco Pierri (Politecnico di Milano) build an MCP server that mirrors a proprietary hardware design tool and benchmark seven locally deployed open-source models on dependency-ordered engineering workflows, isolating which harness choices actually move reliability.

585Agents
Agent Seer: Synthesizing Scenarios from Specification Understanding

Agent Seer: Synthesizing Scenarios from Specification Understanding

Harish Karumuri, Mahesh Vemula, and David Lopes Pegna show that an MCP specification alone carries enough semantic information to synthesize realistic multi-turn agent evaluation scenarios, with no examples, no live tool access, and no domain tuning.

586Agents
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

Kuan-Hao Tseng and colleagues (University of Sydney) build FaulT-Bench, 200 network troubleshooting scenarios across eight topologies that include false fault reports and wrong root-cause claims, then show SADE, ReAct, and Claude Code all collapse when the network is actually healthy.

587Evaluation
Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows

Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows

Maia Kapur and colleagues run a controlled ablation on a production agentic science platform, using protein function characterization as a verifiable task to separate what federation topology, harness type, model choice, and prompt expertise each contribute.

588Reasoning
Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI

Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI

Architect Labs report Redwood, a frontier inference accelerator whose performance model, RTL, UVM environments, formal proofs, firmware and kernels were generated end to end by an AI system in under two weeks from a specification written by two human architects.

589Agents
Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction

Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction

Yu-Lin Tsai and co-authors (NYCU, Berkeley) present Daydreaming, an execution-only attack that reconstructs a hosted multi-file agent skill purely by submitting the ordinary tasks the service exists to perform.

590Agents
Same Model, Different Harness: Different Coding-Agent Results

Same Model, Different Harness: Different Coding-Agent Results

Sydney Lewis holds the model and task fixed and varies only the harness, showing that a coding agent's benchmark number is a property of the model-plus-harness pair rather than of the weights.

591Evaluation
ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

Rui Xie and Lu Chen (EMNLP Findings) argue that screenshot-and-click is the wrong interface for software-operating agents and build ASIL, which exposes applications through structured JSON observations and code-executable semantic actions.

592Agents
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

Yang Xiao and co-authors present PILOT, a supervisor-worker harness that improves a long-horizon agent while the run is still going rather than after it ends.

593Agents
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

Chenhao Wu and co-authors prove a separation result: any safety monitor scoped to a single agent trajectory is provably useless against an attack whose evidence is spread across iterations of an autonomous loop.

594Agents
From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis

From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis

Haiyu Huang, Zhihan Jiang, Michael Lyu and coauthors show that a general agent like Codex or Claude Code now beats purpose-built RCA agents, and argue the remaining gap lives in the harness, which OpsHarness makes self-evolving.

595Agents
Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

Zhongwen Luan, Xiaoyu Zhang, Ming Hu and coauthors ask whether multi-agent repair methods causally fix failures or merely exploit LLM sampling randomness, and build SymTrace to make the distinction measurable.

596Agents
TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

Jiarui Yan, Weiwei Sun, Sijie Li and Yiming Yang at CMU pair 4,465 human Kaggle trajectories with agent runs on the same competitions under one version-level schema, so the ML-development gap can be read as behavior rather than a final score.

597Agents
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Harnesses are hand-built and then frozen, which means one design has to serve deep research, product generation, and long-horizon coding equally well. JIT-Agent is a model whose output is a harness, synthesized per task.

598Agents
Automata from Agent Traces: Failure and Next-Step Prediction

Automata from Agent Traces: Failure and Next-Step Prediction

Seonglae Cho and colleagues at Holistic AI collapse an entire corpus of agent traces into a single compact finite-state machine, then use that FSM as a substrate for both next-step and failure prediction.

599Agents
When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory

When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory

Kazuki Nakayashiki studies what happens when an agent inherits a consolidated memory containing a constraint that has since been withdrawn, and shows that under a two-record verification budget most agents never look at the provenance path.

600Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026