🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Ye, Li, Luo, Yang and colleagues (Xiaomi LLM Core with Peking University, HKU and Renmin) present CodeMidas, an agentic pipeline that builds executable RL environments for coding agents from source code alone, without relying on issues or commits.

265Agents
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Yining She (Carnegie Mellon, work done at Meta) and Lei Lin (Meta) study how to re-evaluate a production analytics agent with tens of thousands of monthly users without rerunning its full benchmark each time, using 574 historical benchmark runs.

266Evaluation
GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

Xinyu Che, Jiaheng Liu and colleagues at Nanjing University build GameLogicBench, 72 gameplay-logic tasks in Godot projects whose rules are checked at every simulation tick, and use it to show where coding agents fail on runtime behavior.

267Agents
When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

Laskar, Fu and colleagues (Dialpad) test whether improving next-turn metrics under gold history predicts better autonomous multi-turn workflow execution, and find that it does not.

268Agents
ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

Zhang, Ding and colleagues (Alibaba Token Hub and Amap) propose ArenaFlow, an RL framework for open-ended agent tasks that turns tournament rankings of trajectories into step-level and skill-level credit.

269Agents
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

Hu, He, Zhou, Tok, Kang and Chaudhuri (UIUC and Microsoft Research) build BI-Bench from real public BI projects and dashboards and show that frontier LLMs answer fewer than half of end-to-end business-intelligence questions correctly; their BI-Agent closes much of the gap.

270Agents
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Siyuan Liu, Yixin Cao and colleagues at Fudan University and the Meituan LongCat Team introduce DENSE, which turns an agent's own execution traces into structured feedback for a second attempt without needing outcome labels, verifiers or expert annotation.

271Agents
Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Hao Fu, Baiting Zhu, Minglei Chen, Yinjie Huang and Shuai Ding (Meta) describe EvoPilot, a human-gated method for running LLM-agent research loops on a production retrieval system, and report a 37-day campaign on the retrieval stack behind Video Deep Dive.

272Retrieval
Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Renkai Ma, Ruyuan Wan and colleagues at Cincinnati, Penn State, Arizona, South Carolina and Florida International analyze 73,093 first-person Reddit posts about using the OpenClaw agent to find which human values users care about when they delegate work.

273Agents
Loopjacking: Hijacking Human-in-the-Loop Approval

Loopjacking: Hijacking Human-in-the-Loop Approval

Arun Kumar (independent) names Loopjacking: a person approves what they see as operation A while the agent runtime uses that approval for a different operation B, and reproduces it in released agent frameworks.

274Agents
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

An, Jang, Kim, Lee, Park and Lee (KRAFTON) release AgentVidBench, a multi-hop video QA benchmark that tests spatial, temporal and causal reasoning in multimodal agents and scores the solution trajectory as well as the final answer.

275Evaluation
Scaling Discovery through Test-Time Communication

Scaling Discovery through Test-Time Communication

Park (UC Berkeley) with Kontonis, Garg, Krishnamurthy and Papailiopoulos (Microsoft Research) show that agents sharing progress through a common directory at test time outperform the same agents working independently on hard problems.

276Agents
Proxifield: Decentralized Multi-Agent Communication through Semantic Proximity

Proxifield: Decentralized Multi-Agent Communication through Semantic Proximity

Pradyumna Tambwekar, Yenchia Feng, Deep Patel and Karime Maamari (Distyl AI) propose Proxifield, a decentralized communication protocol in which each agent decides every round whom to talk to, based on how semantically close their current states are.

277Agents
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

An Alibaba team introduces RecreationWorld, a five-platform environment where a hybrid computer-use agent must inspect a running reference application and build a faithful reimplementation, mixing GUI exploration, coding and visual verification with no fixed workflow.

278Agents
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Du, Yan, Flores and Kadav (Adobe and Brown University) keep a frontier model frozen while it operates professional design software through more than 230 tools, and let an external procedural memory of natural-language skills grow and improve from real user traffic.

279Memory
CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop

CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop

He, Wu, Zhang, Zhao, He and Li (King's College London) present CoLearn, an agentic tutor that keeps an evidence-grounded memory of each learner's mastery and misconceptions and uses it to choose the next question.

280Agents
Symbolic Temporal Supervision of LLM Agents Using Contracts

Symbolic Temporal Supervision of LLM Agents Using Contracts

Yifeng Xiao and Pierluigi Nuzzo present ContrAgent, which writes required agent behaviors as assume-guarantee contracts in LTLf and compiles each to a DFA that both gates tool calls online and grades recorded traces offline.

281Agents
PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

Xinle Yu, Zhen Wang and colleagues at UC San Diego and Johns Hopkins University present PrimeScientist, which treats deciding where to spend a research agent's limited budget as a sequential decision problem.

282Agents
Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making

Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making

Yu Liu, Wenwen Li, Yifan Dou and Guangnan Ye (Fudan University) test whether LLM agents that improve with interaction history in a public goods game are reasoning about other players or extrapolating statistical patterns from past outcomes.

283Agents
Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

Xinshuai Guo, Junjie Wu and colleagues at Tencent Hunyuan and Tsinghua University propose DualViewEval, which compresses expensive agent benchmarks into small task subsets by modeling both final scores and process signals from agent trajectories.

284Agents
Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

Mojtaba Abdolmaleki, Stefanus Jasin and Boyu Wang formulate the choice of how many agentic workflow runs to execute, and of which types, as a portfolio problem that trades extra correct candidates against compute and selection errors.

285Agents
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

Kratika Bhagtani, Kusha Sridhar and colleagues introduce ERPBench, which evaluates screenshot-only computer-use agents on a live, reproducible ERP system and scores each task against values in its database.

286Agents
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

Sehee Kim, Yumin Choi, Minki Kang and Sung Ju Hwang (KAIST, DeepAuto.ai) present EvolveTrade, which treats a trading agent's system prompt as a text policy that a Policy Agent rewrites from decision traces and realized portfolio returns while the backbone LLM stays fixed.

287Agents
Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition

Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition

Dohun Lee and Hyunwoo Park (Seoul National University) measure structural and intent faithfulness of LLM pricing agents in Bertrand competition and find both are unrelated to whether the agents collude.

288Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026