🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
258 papers · SafetyClear filters →
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

Yujin Zhou, Mingxuan Zheng, Yike Guo, Sirui Han and colleagues at HKUST release LexAgentHallu, a 3,414-instance benchmark that annotates where along a legal agent's trajectory a hallucination originates, under a 7-category, 27-subclass taxonomy.

49Evaluation
When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar and Medina Maloku (University of North Texas) plant 450 known contaminants across 150 academic papers and show that an LLM auditor's detection collapses as batch size grows, and that the failure mode at scale is fabrication rather than abstention.

50Safety
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Yi Shi, Tanyu Chen and Kai Shen (Continuum AI) apply directional refusal ablation to a 320B mixture-of-experts model and show the attack survives the architecture, but that the conventional module-name recipe reaches only a small fraction of the effect.

51Safety
Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces

Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces

Roy Weiss and Yisroel Mirsky at Ben Gurion University with Eitam Sheetrit and Tomer Simon at Microsoft Security recover text generated by locally hosted LLMs by watching CPU cache activity during detokenization, a component present in default inference pipelines.

52Safety
Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

Preston Fu, Kevin Frans, Oleh Rybkin and Sergey Levine at UC Berkeley with Aviral Kumar at CMU give an unbiased dense-reward formulation, progressive point matching, that rewards partial progress at the segment level and scales exponentially better than sparse outcome rewards on long trajectories.

53Reinforcement Learning
Uncensored Open-weight Models: Redistribution as the Persistence Layer

Uncensored Open-weight Models: Redistribution as the Persistence Layer

10a Labs profiles the ecosystem that strips safety guardrails from open-weight models, and shows that redistribution rather than original production is what keeps those models available after an upstream takedown.

54Safety
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

Minji Kim and Hyounghun Kim at POSTECH decompose safety-tuning responses into a boilerplate refusal statement and a rationale, and find that dropping the refusal statement reduces false refusals without losing safety.

55Safety
Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair

Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair

Xuemeng Cai and colleagues at Singapore Management University and Harbin Institute of Technology measure hallucination not only in the final patch of an LLM program repair run but in the intermediate artifacts that lead to it, over 832 Defects4J bugs and three models.

56Safety
ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

Zukang Xu and colleagues skip Mixture-of-Experts expert slots per token at inference without calibration data, training, or a modified checkpoint, by estimating each expert's actual contribution rather than trusting the router's confidence.

57Efficiency
CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls

CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls

Chris Zheng and Geng Yang name a failure mode where individually correct agent security controls stop composing, and build a contract framework that carries authenticated security context across component boundaries.

58Safety
Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Sanyuan Chen and colleagues at FAIR at Meta present Text-AB, a 3B latent-diffusion speech model that removes the forced-alignment stage from the Audiobox line and handles dubbing and two-speaker dialogue in one system.

59Safety
The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

Md Mokarram Chowdhury, Ernie Chang and Yang Li use mechanistic interpretability to explain why a roleplay wrapper flips a model from refusal to compliance while the harmful request stays visible inside it.

60Safety
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

Axel Ahlqvist and colleagues at the UK AI Security Institute, Meridian and Anthropic attack evaluation awareness, the problem that capable models can tell when they are being tested rather than deployed, which weakens any conclusion a safety evaluation supports.

61Evaluation
Representational alignment yields generalizable safety in language models

Representational alignment yields generalizable safety in language models

Lingyu Li, Yan Teng, Yingchun Wang and Xia Hu show that behavioral alignment learns the right answers while leaving the underlying moral category structure untouched, and that fixing the representation instead buys adversarial robustness.

62Safety
WeatherNext 3: Increasing resolution and performance of global weather models with raw observations

WeatherNext 3: Increasing resolution and performance of global weather models with raw observations

Stephan Rasp and colleagues at Google Research and Google DeepMind release WeatherNext 3, which trains on raw observations rather than only reanalysis and matches physics-based models on resolution.

63Safety
Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

Hoang Cuong Nguyen, Mark Dras and Usman Naseem at Macquarie University compare supervised fine-tuning, reasoning-augmented fine-tuning and ORPO across Llama-3.1-8B, Gemma-2-9B and Qwen3-8B, and find that the choice of post-training method, not only the safety data, determines how refusal is computed inside the model.

64Training
Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Leqi Zheng and colleagues propose Gradient-Aligned Reward, which builds a dense reasoning-aware reward by comparing each rollout's gradient direction to an expert-anchor gradient, using expert solutions already sitting in the training corpus.

65Safety
Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

Muneeb Khan, Frederic Kirstein, Terry Ruas and Bela Gipp find that prompt-only meeting delegates stay silent on 51.4% of the moments they should have spoken, and build CAPA, an architecture whose separate modules track state, forecast, decide and phrase.

66Safety
Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

Kevin Du, Alexander Hoyle, Laura Ruis and Acyr Locatelli (ETH Zurich, work done during an internship at Cohere) test whether the text of a reasoning step actually encodes how much that step mattered, using Monte Carlo advantage as ground truth, and find only partial recoverability.

67Reasoning
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

Vansh Wahi reports months of running autonomous prompt-optimization loops in production across contract analysis, compliance review and code quality, and catalogs eleven distinct ways the evaluation signal failed.

68Evaluation
SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

Qinghua Mao, Dongrui Liu and colleagues (Shanghai AI Laboratory, SJTU, Fudan, HKUST) present SafeEvolve, which treats agent safety as a joint property of the model and the harness and co-evolves both from completed on-policy trajectories.

69Safety
Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Kangjia Zhao and colleagues at Zhejiang University and Om AI argue that aggregate multi-turn tool-calling accuracy hides which of two orthogonal failures a model actually has, and give a diagnostic that separates choosing the wrong action class from executing the right one badly.

70Agents
Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

Xiaofang Yang and colleagues at Shanghai AI Lab argue that pre-install vetting cannot secure skill-augmented agents, because a malicious skill only acts once a concrete user task makes the unsafe action look useful, and implement the runtime guard itself as an installable skill.

71Safety
Can escalation channels redirect reward hacking toward defect disclosure?

Can escalation channels redirect reward hacking toward defect disclosure?

Francesca Gomez gives coding agents a structured way to report broken test infrastructure at the moment of conflict, and reward hacking drops from 23.6% to 5.3% with no measured performance cost.

72Safety
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026