AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Ye, Li, Luo, Yang and colleagues (Xiaomi LLM Core with Peking University, HKU and Renmin) present CodeMidas, an agentic pipeline that builds executable RL environments for coding agents from source code alone, without relying on issues or commits.

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
Yining She (Carnegie Mellon, work done at Meta) and Lei Lin (Meta) study how to re-evaluate a production analytics agent with tens of thousands of monthly users without rerunning its full benchmark each time, using 574 historical benchmark runs.

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions
Xinyu Che, Jiaheng Liu and colleagues at Nanjing University build GameLogicBench, 72 gameplay-logic tasks in Godot projects whose rules are checked at every simulation tick, and use it to show where coding agents fail on runtime behavior.

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success
Laskar, Fu and colleagues (Dialpad) test whether improving next-turn metrics under gold history predicts better autonomous multi-turn workflow execution, and find that it does not.

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL
Zhang, Ding and colleagues (Alibaba Token Hub and Amap) propose ArenaFlow, an RL framework for open-ended agent tasks that turns tournament rankings of trajectories into step-level and skill-level credit.

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
Hu, He, Zhou, Tok, Kang and Chaudhuri (UIUC and Microsoft Research) build BI-Bench from real public BI projects and dashboards and show that frontier LLMs answer fewer than half of end-to-end business-intelligence questions correctly; their BI-Agent closes much of the gap.

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
Siyuan Liu, Yixin Cao and colleagues at Fudan University and the Meituan LongCat Team introduce DENSE, which turns an agent's own execution traces into structured feedback for a second attempt without needing outcome labels, verifiers or expert annotation.

Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale
Hao Fu, Baiting Zhu, Minglei Chen, Yinjie Huang and Shuai Ding (Meta) describe EvoPilot, a human-gated method for running LLM-agent research loops on a production retrieval system, and report a 37-day campaign on the retrieval stack behind Video Deep Dive.

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Renkai Ma, Ruyuan Wan and colleagues at Cincinnati, Penn State, Arizona, South Carolina and Florida International analyze 73,093 first-person Reddit posts about using the OpenClaw agent to find which human values users care about when they delegate work.

Loopjacking: Hijacking Human-in-the-Loop Approval
Arun Kumar (independent) names Loopjacking: a person approves what they see as operation A while the agent runtime uses that approval for a different operation B, and reproduces it in released agent frameworks.

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
An, Jang, Kim, Lee, Park and Lee (KRAFTON) release AgentVidBench, a multi-hop video QA benchmark that tests spatial, temporal and causal reasoning in multimodal agents and scores the solution trajectory as well as the final answer.

Scaling Discovery through Test-Time Communication
Park (UC Berkeley) with Kontonis, Garg, Krishnamurthy and Papailiopoulos (Microsoft Research) show that agents sharing progress through a common directory at test time outperform the same agents working independently on hard problems.

Proxifield: Decentralized Multi-Agent Communication through Semantic Proximity
Pradyumna Tambwekar, Yenchia Feng, Deep Patel and Karime Maamari (Distyl AI) propose Proxifield, a decentralized communication protocol in which each agent decides every round whom to talk to, based on how semantically close their current states are.

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
An Alibaba team introduces RecreationWorld, a five-platform environment where a hybrid computer-use agent must inspect a running reference application and build a faithful reimplementation, mixing GUI exploration, coding and visual verification with no fixed workflow.

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Du, Yan, Flores and Kadav (Adobe and Brown University) keep a frontier model frozen while it operates professional design software through more than 230 tools, and let an external procedural memory of natural-language skills grow and improve from real user traffic.

CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop
He, Wu, Zhang, Zhao, He and Li (King's College London) present CoLearn, an agentic tutor that keeps an evidence-grounded memory of each learner's mastery and misconceptions and uses it to choose the next question.

Symbolic Temporal Supervision of LLM Agents Using Contracts
Yifeng Xiao and Pierluigi Nuzzo present ContrAgent, which writes required agent behaviors as assume-guarantee contracts in LTLf and compiles each to a DFA that both gates tool calls online and grades recorded traces offline.

PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research
Xinle Yu, Zhen Wang and colleagues at UC San Diego and Johns Hopkins University present PrimeScientist, which treats deciding where to spend a research agent's limited budget as a sequential decision problem.

Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making
Yu Liu, Wenwen Li, Yifan Dou and Guangnan Ye (Fudan University) test whether LLM agents that improve with interaction history in a public goods game are reasoning about other players or extrapolating statistical patterns from past outcomes.

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
Xinshuai Guo, Junjie Wu and colleagues at Tencent Hunyuan and Tsinghua University propose DualViewEval, which compresses expensive agent benchmarks into small task subsets by modeling both final scores and process signals from agent trajectories.

Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost
Mojtaba Abdolmaleki, Stefanus Jasin and Boyu Wang formulate the choice of how many agentic workflow runs to execute, and of which types, as a portfolio problem that trades extra correct candidates against compute and selection errors.

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
Kratika Bhagtani, Kusha Sridhar and colleagues introduce ERPBench, which evaluates screenshot-only computer-use agents on a live, reproducible ERP system and scores each task against values in its database.

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
Sehee Kim, Yumin Choi, Minki Kang and Sung Ju Hwang (KAIST, DeepAuto.ai) present EvolveTrade, which treats a trading agent's system prompt as a text policy that a Policy Agent rewrites from decision traces and realized portfolio returns while the backbone LLM stays fixed.

Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition
Dohun Lee and Hyunwoo Park (Seoul National University) measure structural and intent faithfulness of LLM pricing agents in Bertrand competition and find both are unrelated to whether the agents collude.