🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
258 papers · SafetyClear filters →
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

Albert Wu, Nicholas Roberts and colleagues at the University of Wisconsin-Madison and Princeton wrap an LLM coding agent in a multi-agent pipeline that writes its own formal specifications, so the generated program carries a machine-checkable safety guarantee instead of a test-passing record.

25Agents
Can Data Attribution Filter Out Subliminal Learning? Not Reliably

Can Data Attribution Filter Out Subliminal Learning? Not Reliably

Moritz Weckbecker and co-authors test whether gradient-based training data attribution can find the examples that carry a subliminally transmitted trait, and find that it works inconsistently.

26Safety
Stress-testing Alignment Midtraining

Stress-testing Alignment Midtraining

Sid Baines and colleagues test the assumptions behind alignment midtraining, which continues pretraining on alignment-relevant documents to encourage generalization, at up to 110B parameters and 1B midtraining tokens.

27Safety
Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines

Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines

Yibo Hu at Illinois Institute of Technology identifies verification-status laundering, where an agent handoff keeps the claim that an action was authorised but drops the fact that the claim was never verified, and measures the effect on nine open-weight monitors and two hosted models.

28Safety
Closed-World Resolution Against Tool Hallucination in LLM Agents

Closed-World Resolution Against Tool Hallucination in LLM Agents

Laxmipriya Ganesh Iyer shows that tool-selection and tool-gating defenses cannot address calls to tools that do not exist, gives a five-class taxonomy of tool hallucination, and measures the problem across ten hosted models and the Model Context Protocol.

29Safety
Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

Sarah Radway, Andrew Cheng, Vijay Janapa Reddi and James Mickens at Harvard show that a model can identify which inference engine is executing it from its own output behaviour, then use engine-specific exploits reachable purely through generated tokens.

30Safety
Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

Ashwini Kurady, Sri Sai Charith Grandhi, Rajesh Gupta (RunCtrl) and Sumit Mamoria define Compositional Policy Violations, cases where every step of an agentic workflow passes its own check while the full execution violates an organizational policy, and propose runtime checks over complete execution traces.

31Safety
One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs

One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs

Yibo Hu (Illinois Institute of Technology) shows that filtering harmful peer-induced revisions in multi-agent LLM systems reduces to the model knowing whether its original answer was correct, and that this self-knowledge sets a hard ceiling on any such filter.

32Agents
Do Frontier Models Seek Safety Evidence Before Acting?

Do Frontier Models Seek Safety Evidence Before Acting?

Omer Tafveez (University of Michigan) introduces SAFE, a benchmark that tests whether frontier models choose to retrieve optional safety evidence before making a deployment decision, varying the evidence's retrieval cost, probability, severity and presentation.

33Safety
Latent Undertow: How Ordinary Typos Break Probes

Latent Undertow: How Ordinary Typos Break Probes

Elad David, Max Fomin and Amit LeVi (Zenity) show that ordinary typos, which leave model behavior unchanged, rotate hidden states enough to break activation probes for prompt injection.

34Safety
Position: AI Is Not Ready for Strategic Conflicts

Position: AI Is Not Ready for Strategic Conflicts

Mark Riedl and Glenn Matlin (Georgia Tech) argue that no LLM-enabled wargame should inform planning, doctrine, policy or crisis response without an auditable safety case.

35Safety
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

Aashiq Muhamed, Mona Diab and Virginia Smith (CMU) defend open-weight models against refusal-direction ablation by planting a nonlinear decoy signal that corrupts the attacker's direction estimate.

36Safety
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

Sadia Asif and Mohammad Mohammadi Amiri (RPI) with Prasanna Sattigeri and colleagues (IBM Research) introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents.

37Safety
Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails

Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails

Harish Gaggar (Intuit Credit Karma) compares five context-trimming strategies for multi-step agent workflows and finds that preserving protocol-critical state matters more than the amount of text removed.

38Memory
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs

Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs

Kisu Yang, Yoonna Jang and Heuiseok Lim (Korea University) show that chat templates let any prompt text forge turn and tool-result boundaries, and propose nameless tokenization, which removes the surface strings of control tokens.

39Safety
Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking

Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking

Arun Jose and Julian Stastny (Redwood Research) test whether synthetic document finetuning (SDF) during midtraining can inoculate a model against the broad misalignment that follows from learning to reward hack, and find that it changes what the model says without preventing the misalignment.

40Safety
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Safety evaluations usually ask whether a model refuses a harmful request. Microsoft studies what happens when nobody ever sends that request, and a weaker unaligned model asks for the pieces instead.

41Safety
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Keertana Chidambaram, Andrew Ilyas and Vasilis Syrgkanis (Stanford and CMU) show that planting a harmful but benign-sounding plan in a reasoning model's context makes the model carry out adversarial actions while restating the plan as its own reasoning, which lets it evade chain-of-thought monitors.

42Safety
Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints

Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints

Junlong Shen and Xingyu Li (University of Alberta) audit released batch-normalized unlearning checkpoints and show that published unlearning numbers change when the batch-norm statistics are refit on kept data, with the weights left bit-identical.

43Safety
Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

Narcis Marincat (independent researcher) tests whether independently trained societies of language-model cells that communicate through latent packets share one packet language, and finds that they do not, and that an inherited communication interface can hurt later learning.

44Safety
When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang and JinYeong Bak (Sungkyunkwan University) show that surface noise such as typos makes LLM judges report social bias that is not in the text, so bias measured on noisy text is overestimated.

45Evaluation
Artificial Id: Drive and Persistent Alignment in Agentic AI

Artificial Id: Drive and Persistent Alignment in Agentic AI

Yakov Pyotr Shkolnikov (independent researcher) proposes an artificial id, an internal adaptive drive that decides whether an agent's behavior should continue, stop or change, and argues that alignment must then apply to the continuing system rather than single responses.

46Safety
Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

Jae Gon Kim, Donghoon Yoo and colleagues (Xenoscube) measure NVIDIA's Max-Q inference power profile on a disaggregated B200 serving system and replace it with separate, calibrated power settings for the prefill GPUs and the decode GPUs.

47Efficiency
SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard: Inference-Time Backdoor Detection For Free

Rui Wen (Institute of Science Tokyo), Ahmed Salem and Andrew Paverd (Microsoft Security Response Center), Mark Russinovich (Microsoft Azure) and Zheng Li (Shandong University) detect backdoor activation at inference time by reading the draft-token acceptance rate that speculative decoding already computes, so detection adds no model computation.

48Safety
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026