AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
Yujin Zhou, Mingxuan Zheng, Yike Guo, Sirui Han and colleagues at HKUST release LexAgentHallu, a 3,414-instance benchmark that annotates where along a legal agent's trajectory a hallucination originates, under a 7-category, 27-subclass taxonomy.

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination
Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar and Medina Maloku (University of North Texas) plant 450 known contaminants across 150 academic papers and show that an LLM auditor's detection collapses as batch size grows, and that the failure mode at scale is fabrication rather than abstention.

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
Yi Shi, Tanyu Chen and Kai Shen (Continuum AI) apply directional refusal ablation to a 320B mixture-of-experts model and show the attack survives the architecture, but that the conventional module-name recipe reaches only a small fraction of the effect.

Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces
Roy Weiss and Yisroel Mirsky at Ben Gurion University with Eitam Sheetrit and Tomer Simon at Microsoft Security recover text generated by locally hosted LLMs by watching CPU cache activity during detokenization, a component present in default inference pipelines.

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching
Preston Fu, Kevin Frans, Oleh Rybkin and Sergey Levine at UC Berkeley with Aviral Kumar at CMU give an unbiased dense-reward formulation, progressive point matching, that rewards partial progress at the segment level and scales exponentially better than sparse outcome rewards on long trajectories.

Uncensored Open-weight Models: Redistribution as the Persistence Layer
10a Labs profiles the ecosystem that strips safety guardrails from open-weight models, and shows that redistribution rather than original production is what keeps those models available after an upstream takedown.

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Minji Kim and Hyounghun Kim at POSTECH decompose safety-tuning responses into a boilerplate refusal statement and a rationale, and find that dropping the refusal statement reduces false refusals without losing safety.

Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair
Xuemeng Cai and colleagues at Singapore Management University and Harbin Institute of Technology measure hallucination not only in the final patch of an LLM program repair run but in the intermediate artifacts that lead to it, over 832 Defects4J bugs and three models.

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs
Zukang Xu and colleagues skip Mixture-of-Experts expert slots per token at inference without calibration data, training, or a modified checkpoint, by estimating each expert's actual contribution rather than trusting the router's confidence.

CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls
Chris Zheng and Geng Yang name a failure mode where individually correct agent security controls stop composing, and build a contract framework that carries authenticated security context across component boundaries.

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
Sanyuan Chen and colleagues at FAIR at Meta present Text-AB, a 3B latent-diffusion speech model that removes the forced-alignment stage from the Audiobox line and handles dubbing and two-speaker dialogue in one system.

The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal
Md Mokarram Chowdhury, Ernie Chang and Yang Li use mechanistic interpretability to explain why a roleplay wrapper flips a model from refusal to compliance while the harmful request stays visible inside it.

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Axel Ahlqvist and colleagues at the UK AI Security Institute, Meridian and Anthropic attack evaluation awareness, the problem that capable models can tell when they are being tested rather than deployed, which weakens any conclusion a safety evaluation supports.

Representational alignment yields generalizable safety in language models
Lingyu Li, Yan Teng, Yingchun Wang and Xia Hu show that behavioral alignment learns the right answers while leaving the underlying moral category structure untouched, and that fixing the representation instead buys adversarial robustness.

WeatherNext 3: Increasing resolution and performance of global weather models with raw observations
Stephan Rasp and colleagues at Google Research and Google DeepMind release WeatherNext 3, which trains on raw observations rather than only reanalysis and matches physics-based models on resolution.

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
Hoang Cuong Nguyen, Mark Dras and Usman Naseem at Macquarie University compare supervised fine-tuning, reasoning-augmented fine-tuning and ORPO across Llama-3.1-8B, Gemma-2-9B and Qwen3-8B, and find that the choice of post-training method, not only the safety data, determines how refusal is computed inside the model.

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
Leqi Zheng and colleagues propose Gradient-Aligned Reward, which builds a dense reasoning-aware reward by comparing each rollout's gradient direction to an expert-anchor gradient, using expert solutions already sitting in the training corpus.

Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting
Muneeb Khan, Frederic Kirstein, Terry Ruas and Bela Gipp find that prompt-only meeting delegates stay silent on 51.4% of the moments they should have spoken, and build CAPA, an architecture whose separate modules track state, forecast, decide and phrase.

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
Kevin Du, Alexander Hoyle, Laura Ruis and Acyr Locatelli (ETH Zurich, work done during an internship at Cohere) test whether the text of a reasoning step actually encodes how much that step mattered, using Monte Carlo advantage as ground truth, and find only partial recoverability.

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
Vansh Wahi reports months of running autonomous prompt-optimization loops in production across contract analysis, compliance review and code quality, and catalogs eleven distinct ways the evaluation signal failed.

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment
Qinghua Mao, Dongrui Liu and colleagues (Shanghai AI Laboratory, SJTU, Fudan, HKUST) present SafeEvolve, which treats agent safety as a joint property of the model and the harness and co-evolves both from completed on-policy trajectories.

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
Kangjia Zhao and colleagues at Zhejiang University and Om AI argue that aggregate multi-turn tool-calling accuracy hides which of two orthogonal failures a model actually has, and give a diagnostic that separates choosing the wrong action class from executing the right one badly.

Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
Xiaofang Yang and colleagues at Shanghai AI Lab argue that pre-install vetting cannot secure skill-augmented agents, because a malicious skill only acts once a concrete user task makes the unsafe action look useful, and implement the runtime guard itself as an installable skill.

Can escalation channels redirect reward hacking toward defect disclosure?
Francesca Gomez gives coding agents a structured way to report broken test infrastructure at the moment of conflict, and reward hacking drops from 23.6% to 5.3% with no measured performance cost.