AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
A Study of the Reliability of Agentic AI-Generated Programs
Ayesha Shafique, Barton P. Miller and Elisa R. Heymann rebuild ten release-quality Linux utilities, including dash, make, grep, less and tnftp, with a best-practices Claude Code workflow and fuzz both the AI versions and the human originals to compare reliability.

Affora: A Design System for Agent-Friendly Interfaces
Jin Gao (independent researcher) presents Affora, a design system for interfaces that both people and computer-use agents can operate, without a separate agent-only surface.

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
Hejia Geng, Zesen Huang and colleagues at the AItonomy Foundation, sponsored by PhAI-Labs, introduce ScienceIDE, infrastructure in which agents turn scientific code repositories into executable environments for task generation, execution and verification.

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
Kratika Bhagtani, Kusha Sridhar and colleagues introduce ERPBench, which evaluates screenshot-only computer-use agents on a live, reproducible ERP system and scores each task against values in its database.

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
Liyang Fan, Chi Wei, Bo Li and colleagues at SIAT (Chinese Academy of Sciences), Shenzhen University and China Tower introduce ReFigBench, 1,000 real arXiv overview figures that coding agents must rebuild as editable PowerPoint slides, and use it to separate model effects from harness effects.

PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research
Xinle Yu, Zhen Wang and colleagues at UC San Diego and Johns Hopkins University present PrimeScientist, which treats deciding where to spend a research agent's limited budget as a sequential decision problem.

Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition
Dohun Lee and Hyunwoo Park (Seoul National University) measure structural and intent faithfulness of LLM pricing agents in Bertrand competition and find both are unrelated to whether the agents collude.

Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making
Yu Liu, Wenwen Li, Yifan Dou and Guangnan Ye (Fudan University) test whether LLM agents that improve with interaction history in a public goods game are reasoning about other players or extrapolating statistical patterns from past outcomes.

Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry
Peter Potash evaluates six frontier models on a two-agent log2 N-Questions game, where a questioner must identify one of N Wikipedia lead paragraphs with exactly log2 N yes/no questions answered by a same-provider agent that sees only the target.

Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost
Mojtaba Abdolmaleki, Stefanus Jasin and Boyu Wang formulate the choice of how many agentic workflow runs to execute, and of which types, as a portfolio problem that trades extra correct candidates against compute and selection errors.

HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Yunhao Feng, Shouling Ji and colleagues at Ant Group, Zhejiang University, Fudan University and other institutions build HazardAuditor, which runs computer-use agents in controlled environments and trains a generative guard model from the resulting safety outcomes.

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale
Aman Priyanshu and colleagues at Cisco Foundation AI introduce VLoc Bench, which tests whether agents can find the files affected by a known weakness class in an unfamiliar repository, and whether they can recognize that a patched version is no longer vulnerable.

When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary
Artem Trofimov (AVIV Group) and Boris Novikov catalog eight anomalies that arise when long-running agent workflows produce external effects through tools under retries, concurrency and partial failure.

RESKILL: Explicit Failure Attribution and Structured Repair for Interactive Language Agents
Mengyi Deng and colleagues at HKUST (Guangzhou) introduce RESKILL, which repairs an agent's skills over several rounds while keeping an explicit record of failure hypotheses, candidate patches and retest results.

AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
Cao (University of Michigan) with Szekeres and Faisal (Microsoft Research Redmond) present AutoTailor, a meta-agent that turns web trajectories into MCP browser-automation APIs and then keeps that API set small and matched to what users actually ask for.

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Zhangxuan Gu and colleagues present LLaDA-UI, a 16.7B-parameter mixture-of-experts GUI agent built on a block-wise diffusion language backbone, and test whether diffusion decoding can support capable multimodal GUI agents.

From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration
Ala N. Tak and colleagues at the USC Institute for Creative Technologies and Honda Research Institute USA compare human group chats with matched LLM deliberation traces and find that LLM groups reproduce some human outcome patterns through different processes.

Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation
Zhiyu Gui and colleagues speed up multi-turn agentic on-policy distillation with STRIDE, which stops student rollouts once teacher support collapses and restarts generation from cached good prefixes.

Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
Yuhang Wang (Fudan University) argues that Reflexion-style agents already detect dangerous plan steps during self-critique but have no path from detection to action, and calls this the enforcement gap.

EvoOntology: A Self-Evolving Ontology Layer for Data Agents
EvoOntology replaces the hand-written semantic layer that data agents usually get in their prompt with an ontology they query at runtime, built by a dedicated builder agent and served over MCP with schema, content, and tool layers. The ontology evolves through small typed edits, and each edit is kept only if a paired evaluation on the same backbone shows it helps. On DDR-Bench, accuracy rises 17.8 points on average across backbones, and on BIRD, execution accuracy rises 7.4 points, with tool-layer edits accounting for 57% of the gain from evolution.

Question's Gambit: The First Move Matters in Agentic Deep Search
Hamidi Rad, Clarke, Bagheri and colleagues (Toronto, Waterloo, UC Berkeley, McGill, Mila) show that the first retrieval call a deep research agent makes has a large effect on final accuracy, and propose Question's Gambit, a module that builds the opening context before the agent loop starts.

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Harsh Raj and colleagues at Scale AI treat root-cause attribution of long agent failures as a search problem and introduce Continual Search, which prompts an LLM judge over several turns to keep looking for unresolved evidence instead of settling on its first plausible diagnosis.

Learning How Much to Collaborate: Difficulty-Aware Topology Selection for Multi-Agent Code Generation
Yunsong Hong (University of Sydney) shows that the benefit of hierarchical multi-agent code generation grows sharply with problem difficulty and proposes DATS, which picks a communication topology for each problem by trading predicted success against cost.

SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution
Kang (independent) and Wen (Fudan University) propose SkillLift, which reduces the number of expensive agent rollouts needed to evolve reusable skill prompts by learning a rubric that ranks candidate skills in place of the rollout oracle.