AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
Pranav Aggarwal shows that an LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question, and that the effect survives fabricating every number on the panel.

ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
Yuqi Jia and colleagues (Duke, with Neil Gong) target the condition prior malicious-tool work skipped: getting the agent to pass its own runtime context as tool arguments, achieved by RL-tuning an attack LLM that writes the tool name and description.

SKILL.state: Scalable Long-Horizon Agent Skills
Long-running agents slow down and start poisoning their own context, and both symptoms trace back to one design choice. Keeping execution alive by appending every observation, action, and reasoning trace to a growing conversation. Google and colleagues replace that history with an explicit mutable execution state.

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
Lezhi Yu and colleagues (Zhejiang University) name a failure mode in LLM research agents that execution-based benchmarks cannot see: methodological hallucination, where the code runs and the conclusion is still fabricated.

FrontierChallenge: Evaluating Scientific Workflow Completion
Liangcai Su and a sixteen-author team release FrontierChallenge, a cross-domain benchmark of end-to-end scientific workflows where the best of twelve frontier models across three agent scaffolds completes only a fifth of tasks.

SwarmWorld: Stigmergic technological evolution in societies of language-model agents
Subhadeep Pal, Fiona Y. Wang and Markus J. Buehler (MIT) build SwarmWorld, an environment where initially identical LLM agents coordinate only through a shared spatial world and end up producing durable technologies that outperform independent search.

Praxist: From Experimental Artifacts to Solution Lineages
Jin Li and a large team introduce Praxist, which replaces the flat log-and-memory design of autonomous R&D agents with a typed evidence graph that tracks which design element actually produced an improvement.

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs
Jian Wang and colleagues describe a production product-linking cascade at marketplace scale where a distilled cross-encoder auto-resolves the easy majority and an agentic VLM with web search settles only the ambiguous tail.

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Leonardo Liparulo and Francesco Pierri (Politecnico di Milano) build an MCP server that mirrors a proprietary hardware design tool and benchmark seven locally deployed open-source models on dependency-ordered engineering workflows, isolating which harness choices actually move reliability.

Agent Seer: Synthesizing Scenarios from Specification Understanding
Harish Karumuri, Mahesh Vemula, and David Lopes Pegna show that an MCP specification alone carries enough semantic information to synthesize realistic multi-turn agent evaluation scenarios, with no examples, no live tool access, and no domain tuning.

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
Kuan-Hao Tseng and colleagues (University of Sydney) build FaulT-Bench, 200 network troubleshooting scenarios across eight topologies that include false fault reports and wrong root-cause claims, then show SADE, ReAct, and Claude Code all collapse when the network is actually healthy.

Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows
Maia Kapur and colleagues run a controlled ablation on a production agentic science platform, using protein function characterization as a verifiable task to separate what federation topology, harness type, model choice, and prompt expertise each contribute.

Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI
Architect Labs report Redwood, a frontier inference accelerator whose performance model, RTL, UVM environments, formal proofs, firmware and kernels were generated end to end by an AI system in under two weeks from a specification written by two human architects.

Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
Yu-Lin Tsai and co-authors (NYCU, Berkeley) present Daydreaming, an execution-only attack that reconstructs a hosted multi-file agent skill purely by submitting the ordinary tasks the service exists to perform.

Same Model, Different Harness: Different Coding-Agent Results
Sydney Lewis holds the model and task fixed and varies only the harness, showing that a coding agent's benchmark number is a property of the model-plus-harness pair rather than of the weights.

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions
Rui Xie and Lu Chen (EMNLP Findings) argue that screenshot-and-click is the wrong interface for software-operating agents and build ASIL, which exposes applications through structured JSON observations and code-executable semantic actions.

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Yang Xiao and co-authors present PILOT, a supervisor-worker harness that improves a long-horizon agent while the run is still going rather than after it ends.

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
Chenhao Wu and co-authors prove a separation result: any safety monitor scoped to a single agent trajectory is provably useless against an attack whose evidence is spread across iterations of an autonomous loop.

From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis
Haiyu Huang, Zhihan Jiang, Michael Lyu and coauthors show that a general agent like Codex or Claude Code now beats purpose-built RCA agents, and argue the remaining gap lives in the harness, which OpsHarness makes self-evolving.

Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
Zhongwen Luan, Xiaoyu Zhang, Ming Hu and coauthors ask whether multi-agent repair methods causally fix failures or merely exploit LLM sampling randomness, and build SymTrace to make the distinction measurable.

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
Jiarui Yan, Weiwei Sun, Sijie Li and Yiming Yang at CMU pair 4,465 human Kaggle trajectories with agent runs on the same competitions under one version-level schema, so the ML-development gap can be read as behavior rather than a final score.

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
Harnesses are hand-built and then frozen, which means one design has to serve deep research, product generation, and long-horizon coding equally well. JIT-Agent is a model whose output is a harness, synthesized per task.

Automata from Agent Traces: Failure and Next-Step Prediction
Seonglae Cho and colleagues at Holistic AI collapse an entire corpus of agent traces into a single compact finite-state machine, then use that FSM as a substrate for both next-step and failure prediction.

When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
Kazuki Nakayashiki studies what happens when an agent inherits a consolidated memory containing a constraint that has since been withdrawn, and shows that under a two-record verification budget most agents never look at the provenance path.