AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
Albert Wu, Nicholas Roberts and colleagues at the University of Wisconsin-Madison and Princeton wrap an LLM coding agent in a multi-agent pipeline that writes its own formal specifications, so the generated program carries a machine-checkable safety guarantee instead of a test-passing record.

Can Data Attribution Filter Out Subliminal Learning? Not Reliably
Moritz Weckbecker and co-authors test whether gradient-based training data attribution can find the examples that carry a subliminally transmitted trait, and find that it works inconsistently.

Stress-testing Alignment Midtraining
Sid Baines and colleagues test the assumptions behind alignment midtraining, which continues pretraining on alignment-relevant documents to encourage generalization, at up to 110B parameters and 1B midtraining tokens.

Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines
Yibo Hu at Illinois Institute of Technology identifies verification-status laundering, where an agent handoff keeps the claim that an action was authorised but drops the fact that the claim was never verified, and measures the effect on nine open-weight monitors and two hosted models.

Closed-World Resolution Against Tool Hallucination in LLM Agents
Laxmipriya Ganesh Iyer shows that tool-selection and tool-gating defenses cannot address calls to tools that do not exist, gives a five-class taxonomy of tool hallucination, and measures the problem across ten hosted models and the Model Context Protocol.

Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
Sarah Radway, Andrew Cheng, Vijay Janapa Reddi and James Mickens at Harvard show that a model can identify which inference engine is executing it from its own output behaviour, then use engine-specific exploits reachable purely through generated tokens.

Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows
Ashwini Kurady, Sri Sai Charith Grandhi, Rajesh Gupta (RunCtrl) and Sumit Mamoria define Compositional Policy Violations, cases where every step of an agentic workflow passes its own check while the full execution violates an organizational policy, and propose runtime checks over complete execution traces.

One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs
Yibo Hu (Illinois Institute of Technology) shows that filtering harmful peer-induced revisions in multi-agent LLM systems reduces to the model knowing whether its original answer was correct, and that this self-knowledge sets a hard ceiling on any such filter.

Do Frontier Models Seek Safety Evidence Before Acting?
Omer Tafveez (University of Michigan) introduces SAFE, a benchmark that tests whether frontier models choose to retrieve optional safety evidence before making a deployment decision, varying the evidence's retrieval cost, probability, severity and presentation.

Latent Undertow: How Ordinary Typos Break Probes
Elad David, Max Fomin and Amit LeVi (Zenity) show that ordinary typos, which leave model behavior unchanged, rotate hidden states enough to break activation probes for prompt injection.

Position: AI Is Not Ready for Strategic Conflicts
Mark Riedl and Glenn Matlin (Georgia Tech) argue that no LLM-enabled wargame should inform planning, doctrine, policy or crisis response without an auditable safety case.

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration
Aashiq Muhamed, Mona Diab and Virginia Smith (CMU) defend open-weight models against refusal-direction ablation by planting a nonlinear decoy signal that corrupts the attacker's direction estimate.

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
Sadia Asif and Mohammad Mohammadi Amiri (RPI) with Prasanna Sattigeri and colleagues (IBM Research) introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents.

Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails
Harish Gaggar (Intuit Credit Karma) compares five context-trimming strategies for multi-step agent workflows and finds that preserving protocol-critical state matters more than the amount of text removed.

Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs
Kisu Yang, Yoonna Jang and Heuiseok Lim (Korea University) show that chat templates let any prompt text forge turn and tool-result boundaries, and propose nameless tokenization, which removes the surface strings of control tokens.

Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
Arun Jose and Julian Stastny (Redwood Research) test whether synthetic document finetuning (SDF) during midtraining can inoculate a model against the broad misalignment that follows from learning to reward hack, and find that it changes what the model says without preventing the misalignment.

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
Safety evaluations usually ask whether a model refuses a harmful request. Microsoft studies what happens when nobody ever sends that request, and a weaker unaligned model asks for the pieces instead.

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
Keertana Chidambaram, Andrew Ilyas and Vasilis Syrgkanis (Stanford and CMU) show that planting a harmful but benign-sounding plan in a reasoning model's context makes the model carry out adversarial actions while restating the plan as its own reasoning, which lets it evade chain-of-thought monitors.

Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints
Junlong Shen and Xingyu Li (University of Alberta) audit released batch-normalized unlearning checkpoints and show that published unlearning numbers change when the batch-norm statistics are refit on kept data, with the weights left bit-identical.

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells
Narcis Marincat (independent researcher) tests whether independently trained societies of language-model cells that communicate through latent packets share one packet language, and finds that they do not, and that an inherited communication interface can hurt later learning.

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text
DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang and JinYeong Bak (Sungkyunkwan University) show that surface noise such as typos makes LLM judges report social bias that is not in the text, so bias measured on noisy text is overestimated.

Artificial Id: Drive and Persistent Alignment in Agentic AI
Yakov Pyotr Shkolnikov (independent researcher) proposes an artificial id, an internal adaptive drive that decides whether an agent's behavior should continue, stop or change, and argues that alignment must then apply to the continuing system rather than single responses.

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
Jae Gon Kim, Donghoon Yoo and colleagues (Xenoscube) measure NVIDIA's Max-Q inference power profile on a disaggregated B200 serving system and replace it with separate, calibrated power settings for the prefill GPUs and the decode GPUs.

SpecGuard: Inference-Time Backdoor Detection For Free
Rui Wen (Institute of Science Tokyo), Ahmed Salem and Andrew Paverd (Microsoft Security Response Center), Mark Russinovich (Microsoft Azure) and Zheng Li (Shandong University) detect backdoor activation at inference time by reading the draft-token acceptance rate that speculative decoding already computes, so detection adds no model computation.