AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
Zhuowen Liu at the Japan Advanced Institute of Science and Technology re-evaluates fifteen prompt-injection detectors and two task-aware judges by replaying the ground-truth tool calls of AgentDojo and tau-bench, and finds that public benchmark scores do not predict detector behavior inside agents.

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine
Eray Turkel and colleagues at Roblox release OpenGameEval, a benchmark that runs language-model agents inside reproducible Roblox Studio sessions and separates observation tools from editing tools so exploration behavior can be measured directly.

Trained Agentic Context Management
Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
Ankur Samanta, Kaveh Hassani and Anirudh Goyal at Meta AI, with Yonathan Efroni (Tel Aviv) and Paul Sajda (Columbia), introduce MIRA, a research-agent architecture in which an outer meta-reasoner decides what to investigate next and a fresh executor carries out each investigation, and train that outer policy with RL.

Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery
Kaikai Zhang, Dongdong She and colleagues at HKUST show that LLM vulnerability-discovery agents can form hypotheses cheaply but must spend heavily to verify each one, and build RedHerring, a defense that plants provably safe decoy vulnerabilities to absorb that verification budget.

Can Agents Design Libraries for Agents?
Gabriel Orlanski, Frederic Sala and Aws Albarghouthi (UW-Madison), Alex L. Zhang (MIT), Vincent Sunn Chen (Snorkel AI) and Ludwig Schmidt (Stanford) introduce LibraryDesignBench, which scores a library written by one agent only by how well other agents can use it.

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Ashish Jain (Sarvam AI) and Armaan Sandhu (UMass Amherst) introduce UserProxyBench, which scores the simulated user in tau-bench-style agent benchmarks on whether it followed its private instructions, separately from whether the agent succeeded.

Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions
Xunjian Yin, Bhuwan Dhingra, Shuyan Zhou and colleagues at Duke with Amazon collaborators build BreakingWeb, which makes browser tasks harder by changing the environment under a task agents already solve while keeping the instruction and success criterion fixed.

LongCat-DeepResearch Technical Report
The Meituan LongCat Team describes LongCat-DeepResearch, a deep research system that pairs an enhanced LongCat model with a multi-agent workflow built around a compact research plan, ResearchSpec, instead of repeated full-report rewrites.

Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
Rohith Reddy Bellibatlu, Zichong Wang and Wenbin Zhang at Florida International University audit tool-using agent benchmarks by treating each tool's documented interface as an executable contract and checking whether the implementation, and the grader that trusts it, actually do what the interface says.

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi and Leshem Choshen at IBM Research and the Weizmann Institute study which unrequested information an agent should go after, and train an 8B questioner to choose the questions that retrieve the most required evidence.

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
Youling Huang, Lin Lin and colleagues from Kuaishou with DUT, XJTU, Tsinghua and other universities show that on-policy distillation helps agentic RL only while the teacher is ahead of the student, and propose GATS, which scales the distillation term by the measured teacher-student gap and drops it at crossover.

Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety
Charlie Summers, Oliver Kennedy (Buffalo), Eugene Wu and colleagues at Columbia propose Environment Steering: model the agent's and harness's execution state as database tables, track record-level data flow, and when a declarative policy is violated, send the agent specific feedback so it can recover safely.

SAGE: A Statistical Acceptance Gate for Self-Evolving Agents
Yihao Wang (Peking University), with collaborators at Tencent, Imperial College London and Michigan, show that the usual acceptance rule in skill self-evolution, keep any edit that raises the average validation score, admits regressions and noise, and replace it with a paired statistical gate called SAGE.

Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents
Kehang Zhu and David C. Parkes (Harvard) and Anand Shah (MIT) test whether interface designs that make strategic choices easier for humans also improve LLM agents in auctions and matching markets, where the optimal strategy is known.

VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses
Jiexing Qi and colleagues at Huawei's ICT AI Competence Center propose VACE, which alternates agentic RL on the model with trajectory-driven revisions to the harness, and accepts a harness revision only if it improves validation performance with the newly trained model.

Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations
Guanghui Min, Chen Chen and colleagues at the University of Virginia with Nokia study how repeated context compression hurts long-horizon agents and introduce PAIR, which locates the individual compressions that cause failures and rewrites the compression prompt around them.

Rational Clarification by Assistive Agents via Value-of-Information Reasoning
T. Duy Nguyen-Hien and Wee Sun Lee (NUS), Yee Whye Teh (Oxford) and Tan Zhi-Xuan introduce REVOIR, an inference-time method that decides whether an assistant should ask a clarifying question by estimating how much the answer would raise expected task reward, net of the cost of asking.

Follow the Entities: A Corpus Map for Agentic Search
Soyeong Jeong (KAIST, Microsoft intern), Sujay Kumar Jauhar, Sung Ju Hwang and Andrew Nam at Microsoft and KAIST build CorpusMap, an offline entity layer over a document collection that lets a search agent follow shared entities from one document to related ones instead of re-searching the flat corpus.

SelfSearch: Reward-Free Search for Self-Improving Agents
Jungwoo Yang, Injin Kong and Yohan Jo at Seoul National University introduce SelfSearch, in which coding agents rewrite their own instructions, tools and procedures using records of earlier self-modification episodes, with no downstream task reward during the search.

Agents as Software: A Programming Languages Agenda for Agent Reliability
Shraddha Barke (Microsoft Research) and Adithya Murali (UW-Madison) argue in an Onward! 2026 essay that agents should be treated as programs whose behavior is spread across prompts, tools, memories and traces, and lay out how specifications, static analysis and runtime monitoring from programming languages research apply to them.

"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused
Xutao Mao, Rui Qian and colleagues at City University of Hong Kong and Fudan introduce CAVE-Bench, 365 agentic tasks that test whether an agent keeps verified correct work when a later message falsely blames it for a failure, and find that agents damage correct work in up to 60% of runs.

Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
Ritul Satish, Prasoon Sinha, Akiho Kawada and Neeraja Yadwadkar at UT Austin run nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench to separate the three decisions in a context compression policy (mechanism, trigger, and amount removed) and measure how each affects success, tokens, latency and cost.

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
Janvijay Singh (UIUC, Microsoft Research intern), Vaishnavi Shrivastava, Dilek Hakkani-Tür, Ece Kamar and Asli Celikyilmaz at Microsoft Research AI Frontiers introduce AGNI, a pipeline that breaks one environmental assumption behind a successful terminal-agent trajectory while keeping the task solvable, and use it to measure how well agents adapt.