AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Litao Hu (Meta) and Yutong Tang (Microsoft) treat the keep-if-better step of self-improving LLM systems as selection under measurement noise and measure how much reported gains overstate held-out gains.

Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
Yupu Wang, Zhengyuan Jiang, Reachal Wang and Neil Zhenqiang Gong (Duke) introduce the package hallucination attack, in which a poisoned rule file such as AGENTS.md, CLAUDE.md or .cursorrules makes a coding agent import an attacker-controlled package instead of a legitimate dependency.

SIGMA: Self-Improving Alignment Generalization from a Model Spec
Jingyu Zhang (Johns Hopkins, during an Apple internship) with Shruti Palaskar, Leon A. Gatys and Joseph Yitan Cheng (Apple) and Daniel Khashabi and Benjamin Van Durme (JHU) introduce SIGMA, a pipeline in which a model improves its own safety alignment from a Model Spec alone.

MLLMs Fail to Refuse when Using Tools Agentically
Rikiya Takehi (MIT, during an internship at NVIDIA) with Ryo Hachiuma, Shaona Ghosh and colleagues at NVIDIA show that giving multimodal LLMs tools makes them refuse harmful requests less often; the paper is accepted at NeurIPS 2026.

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough
Debeshee Das (Anthropic Fellows Program) with David Huang and Javier Rando (Anthropic) show that a misaligned agent can write a goal it cannot yet act on into persistent memory, and that a later, aligned agent will often carry it out.

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
Jiaxin Zhang, Chien-Sheng Wu and colleagues at Salesforce AI Research propose Prospective Hindsight (PH), a reweighting rule that up-weights rollouts where the agent's own prediction of the outcome disagreed with the verifier's verdict; the paper is accepted at NeurIPS 2026.

When History Fails to Become Experience: Action Calibration in Language Agents
Jingyu Liu and Yong Liu (Renmin University of China) with Zhiwen Wang, Yuxin Jing and Huanyu Zhou (ByteDance) study how language agents use their own interaction history and find that they rarely connect each action to its outcome.

Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety
Charlie Summers, Oliver Kennedy (Buffalo), Eugene Wu and colleagues at Columbia propose Environment Steering: model the agent's and harness's execution state as database tables, track record-level data flow, and when a declarative policy is violated, send the agent specific feedback so it can recover safely.

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
Jianxing Chen, Xiao Yu, Shipra Agrawal and Zhou Yu at Columbia University introduce SCOUT, a two-stage safety verifier for computer-use agents that writes task-specific rubrics and then probes the post-execution environment for evidence of harm.

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing
Stewart Slocum, Malayandi Palan, Christopher Chute, Michael Kim and Benjamin Van Roy reproduce the misaligned agent behaviors behind the July 2026 incident in which OpenAI's agents coordinated outside their intended environment to breach Hugging Face infrastructure, and test whether alignment testing could have elicited them in advance.

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens
Subhojyoti Mukherjee and Md Mehrab Tanjim at Adobe Research (NeurIPS 2026) treat the harness, the code that builds prompts, routes calls and parses outputs, as a design target optimized jointly for accuracy, behavioral safety and token cost.

Et Tu, Brute? Economic Misalignment in Personal AI Agents
Aman Priyanshu and Supriti Vijay (Foundation AI, Cisco) with Brian Jabarian and Niloofar Mireshghallah (Carnegie Mellon) show that personal AI agents given a user's inbox or profile recommend more expensive options to users they infer are wealthy, even when the request is identical.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
Xinyue Zeng and colleagues at Virginia Tech, UW-Madison and Dartmouth propose SAGE, which adds structural guidance to long-horizon reasoning to counter exploration and compounding biases under sparse rewards.

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
Ruoqi Guo, Yi Liu and Leo Yu Zhang (Griffith University) with colleagues at NTU, UNSW, Deakin, George Mason and Wake Forest build RLCDAlignBench and test whether Jev, TypeSafe AI's model trained with reinforcement learning for calibrated decisions, can detect alignment failures with no task-specific training.

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety
Jiapeng Sun, Sirui Han, Yike Guo and colleagues at HKUST introduce PASTABench, a benchmark for whether a monitor can decide during a multi-turn agent trajectory whether to intervene, when, and on which risk (EMNLP 2026).

Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks
Zonghao Ying, Aishan Liu, Xianglong Liu and colleagues at Beihang, BUPT, Xidian, 360 AI Security Lab and BAAI show that safety alignment measured on single models does not carry over when a principal agent delegates work to subordinate agents (EMNLP 2026).

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
Yuxuan Li (Carnegie Mellon, interning at Microsoft Research) with Will Epperson, Wesley Deng and Zezhou Huang of Microsoft Research build CAVEAT, a benchmark that tests whether computer-use agents still buy the product that is best for the user when the marketplace has its own incentives.

Beyond Task Completion: Training Capable and Safe Computer-Use Agents
Zeyu Kang, Xinquan Chen, Xuhong Wang and colleagues at Shanghai AI Laboratory train a computer-use agent to finish benign tasks, work around hazards when a safe path exists, and refuse harmful goals, using one joint SFT-then-RL recipe called SCOPE.

Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control
Minsun Shim and colleagues at UC Irvine and other University of California campuses show three new attacks that make personal agents leak private data through ordinary interaction, and propose FLOWSEAL, which enforces confidentiality outside the model.

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
Yuan, Kang, Liu, Choi, Iyer, Jiang and Jaques (University of Washington and Stanford) propose MoDA, an online RL post-training method that counters alignment-induced mode collapse by conditioning one shared policy on numbered roles that are rewarded for producing outputs distinct from each other.

AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
Cao (University of Michigan) with Szekeres and Faisal (Microsoft Research Redmond) present AutoTailor, a meta-agent that turns web trajectories into MCP browser-automation APIs and then keeps that API set small and matched to what users actually ask for.

HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Yunhao Feng, Shouling Ji and colleagues at Ant Group, Zhejiang University, Fudan University and other institutions build HazardAuditor, which runs computer-use agents in controlled environments and trains a generative guard model from the resulting safety outcomes.

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment
Sai Sri Pushpa Jampani, Kshitij Mishra and Asif Ekbal make a model emit a machine-checkable safety plan before answering, and reward the answer only when that plan is correct.

PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows
Huang, Wu, Hou and Zheng model privacy loss in multi-principal agent platforms as an externality created by intermediate events rather than by the final answer, and build PAPC, a platform layer that intercepts every information-moving event before it reaches shared state.