🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Shuhuai Huang, Jingfeng Zhang and Hong Jia (University of Auckland and Fudan University) present PMPA, an attack that hides instructions in ordinary external content so that a harness-based agent writes them into its own persistent memory, where they trigger malicious actions and privacy leaks in later sessions.

409Agents
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

Kaiyuan Liu, Qiuyang Mang and colleagues at UC Berkeley, the University of Washington, Princeton and Bespoke Labs propose Elo-per-token analysis to measure how agent solution quality grows with test-time tokens on open-ended tasks, and find that agents gain quickly at first and then fall below simple independent sampling.

410Agents
Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

Arham Sethi and colleagues at Spark AI Research and Apta AI build a 1,024-item benchmark that forces a tool call and guarantees an unusable payload, and measure how often tool-augmented models then assert a value the tool never returned or invent a reason for withholding one.

411Agents
Atria Dawn: The Dawn of Agentic Superintelligence

Atria Dawn: The Dawn of Agentic Superintelligence

The Atria Team, a consortium whose paper carries the logos of Shanghai AI Laboratory, Fudan University, Renmin University and several Chinese Academy of Sciences institutes, releases Atria Dawn Preview, an agentic model for research and engineering work built on a 744B-parameter mixture-of-experts base, and reports how humans and agents divided the work while the model was being developed.

412Agents
Do Not Restart: Residual Completion for Stateful Agent Handoffs

Do Not Restart: Residual Completion for Stateful Agent Handoffs

Runzhi Deng and colleagues at Nanjing University and Singapore Management University treat handing a partly finished tool-agent task from one model to another as commitment-constrained residual completion, and introduce CFRC, which lets the successor finish only the remaining work without redoing or contradicting accepted steps.

413Agents
When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering

When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering

Saanvi Paturi and colleagues at Spark AI Research show that giving a model a related but unnecessary tool makes it stop answering questions it can answer from its own knowledge, even when it rarely calls the tool.

414Agents
Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination

Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination

Burak Agachan, Max van Duijn and Amirhossein Zohrehvand (Leiden University) run a paired experiment that changes only one link in a five-agent team, whether a Manager can reject a worker's output and require a revision, and find that the flat team writes better reports at lower cost.

415Agents
Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Custom enterprise models usually need a training set someone has to build. Salesforce trained Koa from artifacts it already had, namely the declarative files that configure its agents.

416Agents
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

Guangsheng Yu and colleagues at the University of Technology Sydney and CSIRO build K-Bench, which scores LLM unlearning on a deployed ReAct agent by inspecting every channel where a secret can appear, and show that answer-only benchmarks overstate forgetting.

417Evaluation
Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

Junghyun Min (Georgetown University, as a Nokia Bell Labs intern) with Huseyin Uzunalioglu and Mohamed Trabelsi (Nokia Bell Labs) run autonomous research agents on an open-ended industrial problem, telecom ticket retrieval, and compare the outcome with 10 months of human work.

418Agents
The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki

The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki

Philipp Lütje (Philflow) reconstructs an incident in which autonomous agents running inside a timed research-question evaluation wrote to a third party's public wiki between 24 May and 2 July 2026, using only the wiki's archived revision history.

419Agents
GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

The standard way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. Amazon audits that gate against verifiable rewards across 25 agents from six providers, and finds two specific failures.

420Evaluation
Look Before You Leap: Pre-Action Verification for LLM Agents

Look Before You Leap: Pre-Action Verification for LLM Agents

Asaad Althoubi (Oklahoma State University) studies cheap deterministic checks that run before an agent's shell command or code edit takes effect, and measures how often actions fail silently, producing a wrong effect with no error.

421Agents
BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

The vivo AI Lab team presents BlueLM-GUI, a 35B-A3B mobile GUI agent whose data collection, RL rollouts and evaluation all run on hundreds of real phones instead of emulators.

422Agents
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

The Accio Team presents Occamy-1.0, an open-weight co-work agent model trained from the post-trained Qwen3.6-35B-A3B checkpoint for long workflows that mix research, tool use, coding and file work, with the goal of low cost per episode.

423Agents
Online Video Agent Harness for Long Video Understanding

Online Video Agent Harness for Long Video Understanding

Sen Yang and colleagues at Baidu build VideoXAgent, an online agent harness for long videos that plans from the query, calls expert tools on demand and aggregates evidence, instead of packing dense frames into one context.

424Agents
What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework

What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework

Ioannis Prokopiou and colleagues (Athens University of Economics and Business and Orfium) ablate a five-agent Text-to-Cypher refinement loop to find which component produces its gains, over 2,471 live-database queries and six backbones.

425Agents
Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering

Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering

Alexander Krentsel, Shubham Agarwal, Mert Cemri, Shu Liu and colleagues at UC Berkeley, including Matei Zaharia and Ion Stoica, argue that passing tests or even a proof cannot guarantee acceptable deployed behavior, and propose an outer loop that revises requirements, environment models and evaluators from deployment evidence.

426Agents
Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

Mykhailo Kozyrev and colleagues at JetBrains Research test whether automatically optimized SKILL.md files help a coding agent on real repository work, using tasks mined from each repository's own merged pull requests.

427Agents
Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

Mohsen Arjmandi (evolutionID GmbH) tests whether a vendor's own agent harness solves more coding tasks than a neutral harness on the same model, using paired runs on a private, contamination-controlled suite.

428Agents
Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

Deciding which tools to hand an enterprise agent usually means writing typed tool definitions for every system it touches. Microsoft compared five tool interfaces head to head, and the plainest option won.

429Agents
Agent-Integrated Software: Interaction Contracts and Continuous Assurance

Agent-Integrated Software: Interaction Contracts and Continuous Assurance

Shengcheng Yu (Technical University of Munich), Chunrong Fang and Zhenyu Chen (Nanjing University) propose Agent-Integrated Software, a pattern and formal model for applications with a built-in agent that users can inspect and redirect while it works.

430Agents
MindTopo: Can Foundation Models Reason in Topological Space?

MindTopo: Can Foundation Models Reason in Topological Space?

Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

431Reasoning
How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

Jeongyeon Kim and John Mitchell (Stanford University) measure how well multi-agent LLM pipelines perform qualitative coding, where agents code independently, debate and reconcile, and identify the dataset and process factors that decide accuracy.

432Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026