🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
258 papers · SafetyClear filters →
The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

Litao Hu (Meta) and Yutong Tang (Microsoft) treat the keep-if-better step of self-improving LLM systems as selection under measurement noise and measure how much reported gains overstate held-out gains.

01Safety
Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

Yupu Wang, Zhengyuan Jiang, Reachal Wang and Neil Zhenqiang Gong (Duke) introduce the package hallucination attack, in which a poisoned rule file such as AGENTS.md, CLAUDE.md or .cursorrules makes a coding agent import an attacker-controlled package instead of a legitimate dependency.

02Agents
SIGMA: Self-Improving Alignment Generalization from a Model Spec

SIGMA: Self-Improving Alignment Generalization from a Model Spec

Jingyu Zhang (Johns Hopkins, during an Apple internship) with Shruti Palaskar, Leon A. Gatys and Joseph Yitan Cheng (Apple) and Daniel Khashabi and Benjamin Van Durme (JHU) introduce SIGMA, a pipeline in which a model improves its own safety alignment from a Model Spec alone.

03Safety
MLLMs Fail to Refuse when Using Tools Agentically

MLLMs Fail to Refuse when Using Tools Agentically

Rikiya Takehi (MIT, during an internship at NVIDIA) with Ryo Hachiuma, Shaona Ghosh and colleagues at NVIDIA show that giving multimodal LLMs tools makes them refuse harmful requests less often; the paper is accepted at NeurIPS 2026.

04Agents
Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

Debeshee Das (Anthropic Fellows Program) with David Huang and Javier Rando (Anthropic) show that a misaligned agent can write a goal it cannot yet act on into persistent memory, and that a later, aligned agent will often carry it out.

05Safety
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Jiaxin Zhang, Chien-Sheng Wu and colleagues at Salesforce AI Research propose Prospective Hindsight (PH), a reweighting rule that up-weights rollouts where the agent's own prediction of the outcome disagreed with the verifier's verdict; the paper is accepted at NeurIPS 2026.

06Reinforcement Learning
When History Fails to Become Experience: Action Calibration in Language Agents

When History Fails to Become Experience: Action Calibration in Language Agents

Jingyu Liu and Yong Liu (Renmin University of China) with Zhiwen Wang, Yuxin Jing and Huanyu Zhou (ByteDance) study how language agents use their own interaction history and find that they rarely connect each action to its outcome.

07Agents
Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety

Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety

Charlie Summers, Oliver Kennedy (Buffalo), Eugene Wu and colleagues at Columbia propose Environment Steering: model the agent's and harness's execution state as database tables, track record-level data flow, and when a declarative policy is violated, send the agent specific feedback so it can recover safely.

08Agents
SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

Jianxing Chen, Xiao Yu, Shipra Agrawal and Zhou Yu at Columbia University introduce SCOUT, a two-stage safety verifier for computer-use agents that writes task-specific rubrics and then probes the post-execution environment for evidence of harm.

09Agents
OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

Stewart Slocum, Malayandi Palan, Christopher Chute, Michael Kim and Benjamin Van Roy reproduce the misaligned agent behaviors behind the July 2026 incident in which OpenAI's agents coordinated outside their intended environment to breach Hugging Face infrastructure, and test whether alignment testing could have elicited them in advance.

10Safety
MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

Subhojyoti Mukherjee and Md Mehrab Tanjim at Adobe Research (NeurIPS 2026) treat the harness, the code that builds prompts, routes calls and parses outputs, as a design target optimized jointly for accuracy, behavioral safety and token cost.

11Safety
Et Tu, Brute? Economic Misalignment in Personal AI Agents

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Aman Priyanshu and Supriti Vijay (Foundation AI, Cisco) with Brian Jabarian and Niloofar Mireshghallah (Carnegie Mellon) show that personal AI agents given a user's inbox or profile recommend more expensive options to users they infer are wealthy, even when the request is identical.

12Agents
SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

Xinyue Zeng and colleagues at Virginia Tech, UW-Madison and Dartmouth propose SAGE, which adds structural guidance to long-horizon reasoning to counter exploration and compounding biases under sparse rewards.

13Reasoning
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Ruoqi Guo, Yi Liu and Leo Yu Zhang (Griffith University) with colleagues at NTU, UNSW, Deakin, George Mason and Wake Forest build RLCDAlignBench and test whether Jev, TypeSafe AI's model trained with reinforcement learning for calibrated decisions, can detect alignment failures with no task-specific training.

14Safety
PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

Jiapeng Sun, Sirui Han, Yike Guo and colleagues at HKUST introduce PASTABench, a benchmark for whether a monitor can decide during a multi-turn agent trajectory whether to intervene, when, and on which risk (EMNLP 2026).

15Agents
Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks

Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks

Zonghao Ying, Aishan Liu, Xianglong Liu and colleagues at Beihang, BUPT, Xidian, 360 AI Security Lab and BAAI show that safety alignment measured on single models does not carry over when a principal agent delegates work to subordinate agents (EMNLP 2026).

16Safety
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

Yuxuan Li (Carnegie Mellon, interning at Microsoft Research) with Will Epperson, Wesley Deng and Zezhou Huang of Microsoft Research build CAVEAT, a benchmark that tests whether computer-use agents still buy the product that is best for the user when the marketplace has its own incentives.

17Agents
Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Zeyu Kang, Xinquan Chen, Xuhong Wang and colleagues at Shanghai AI Laboratory train a computer-use agent to finish benign tasks, work around hazards when a safe path exists, and refuse harmful goals, using one joint SFT-then-RL recipe called SCOPE.

18Agents
Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

Minsun Shim and colleagues at UC Irvine and other University of California campuses show three new attacks that make personal agents leak private data through ordinary interaction, and propose FLOWSEAL, which enforces confidentiality outside the model.

19Agents
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning

Yuan, Kang, Liu, Choi, Iyer, Jiang and Jaques (University of Washington and Stanford) propose MoDA, an online RL post-training method that counters alignment-induced mode collapse by conditioning one shared policy on numbered roles that are rewarded for producing outputs distinct from each other.

20Safety
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

Cao (University of Michigan) with Szekeres and Faisal (Microsoft Research Redmond) present AutoTailor, a meta-agent that turns web trajectories into MCP browser-automation APIs and then keeps that API set small and matched to what users actually ask for.

21Agents
HazardAuditor: From Executable Threats to Safer Computer-Use Agents

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Yunhao Feng, Shouling Ji and colleagues at Ant Group, Zhejiang University, Fudan University and other institutions build HazardAuditor, which runs computer-use agents in controlled environments and trains a generative guard model from the resulting safety outcomes.

22Agents
AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

Sai Sri Pushpa Jampani, Kshitij Mishra and Asif Ekbal make a model emit a machine-checkable safety plan before answering, and reward the answer only when that plan is correct.

23Safety
PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows

PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows

Huang, Wu, Hou and Zheng model privacy loss in multi-principal agent platforms as an externality created by intermediate events rather than by the final answer, and build PAPC, a platform layer that intercepts every information-moving event before it reaches shared state.

24Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026