AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs
Ansuman Mullick and Eray Tuzun (Bilkent University) classify personal facts into a behavioral ontology and apply category-specific retention policies as deterministic functions over LLM-extracted metadata, then locate through ablation which half of the design produces which gain.

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
Xi Shi and Qian Lou (University of Central Florida) build KVShareArena, a benchmark for reusing KV caches when the reused text is not a prompt prefix, which is the case for retrieval-augmented servers and for multi-agent coordinators reading reports written by other agents.

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents
Benjamin Gruenbaum, Doron Porat, Assaf Natanzon and colleagues at Eon generate a complete fictional company, including simulators of Salesforce, Zendesk, Slack and Gong, so enterprise agent answers can be graded exactly against computed answer keys.

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Pujun Zheng at East China Normal University with Shanghai Artificial Intelligence Laboratory audits SWE-Bench Pro, finds reward hacking through gold-solution leakage and task-quality defects, and releases SWE-Bench Pro Verified, on which several models score substantially lower than previously reported.

SkillAdam: Stable and Efficient Skill Evolution for Agents
Gaoyuan Li, Meihao Fan, Shaolei Zhang and Ju Fan at Renmin University present SkillAdam, which ports Adam's two moment estimates to the optimization of discrete, non-differentiable skill documents so that skill self-evolution stops oscillating.

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems
Maximilian Schall, Sedigheh Eslami, Antoine Chaffin and colleagues at Perplexity AI release Q2D-Web, a 190M-document web corpus with 70k agent-reformulated search queries in ten languages, built because production RAG retrievers serve machine-written queries and existing benchmarks test human-written ones.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Sihan Ge and colleagues at Cardinal Operations and Shanghai Jiao Tong University benchmark whether an LLM agent knows when to ask a clarifying question before turning a natural-language operations research request into a mathematical program.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Minji Kim, Jihyoung Jang and Hyounghun Kim at POSTECH argue that non-compliance in vision-language models is evaluated at the wrong granularity, and build a benchmark where a single query mixes answerable content with content that should be withheld.

Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory
Peng Cui, Heejin Do and Mrinmaya Sachan at ETH Zurich apply Knowledge Space Theory, which formalizes the idea that mastering a concept requires mastering its prerequisites, as a normative standard for LLM mathematical knowledge, and compare eight models against real human learners.

Single-Query Black-Box Calibration Auditing via Logit Bias
Roman Plaud and colleagues at Institut Polytechnique de Paris, Onepoint and Ghent University show that a logit_bias parameter is enough to recover exact probability thresholds from an API that hides output probabilities, using one query per sample.

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
Daan Henselmans, Derck Prinzhorn and Arno Libert at the Aithos Research Foundation score how well a model can defend its verdict under critical questioning, using a standard that does not require ground truth about the right answer.

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Shenxi Wu and colleagues at The Chinese University of Hong Kong and Shanghai AI Laboratory build a benchmark that tests scientific document understanding as a research-assistant workflow rather than as isolated perception, retrieval and reasoning tasks, and ship the training data to improve it.

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
Matthias Busch and colleagues at Helmholtz-Zentrum Hereon and Hamburg University of Technology audit 22 frontier models on 12 molecular regression benchmarks to separate models that predict a property from models that reproduce a published number.

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
Abhishek Sharma builds an executable benchmark for agents resolving payment exceptions when a merchant's processor, ledger, ERP and bank feed hold contradictory beliefs about the same order, and grades on executed monetary effects rather than answer accuracy.

RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents
Aziz Ben Amor and colleagues at Pi School release RefactorPlatform, an evaluation harness that holds the environment fixed and varies one coding-agent design axis at a time on 100 repository-scale RefactorBench tasks.

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
Zhibo Yang and colleagues build a benchmark for scientific discovery rather than reproduction: agents see a neutral objective and frozen data, with the source study's conclusions, expected values, and analysis path withheld.

$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction
Quan Shi, Keshav Dhandhania, Karthik Narasimhan and Victor Barres at Sierra and Princeton make agent construction itself the benchmark task: a developer agent must deliver a working customer-service agent under the conditions of a real client engagement.

Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency
Sverrir Thorgeirsson, Theo B. Weidmann and Zhendong Su at ETH Zurich run a preregistered cross-sectional study of 100 tertiary-level students to find out which measured skills predict how well someone performs at vibe coding.

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Ji Soo Lee and colleagues at Meta and KAIST build WearableQA from the wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements each.

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Xinran Zhang and colleagues at GAIR ask whether enterprise-agent rankings survive a change in who the agent is competing against, and find they largely do not.

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models
Ross Tieman and Evan Markou argue that semantic similarity is the wrong diversity measure for populations of language models, and use compression distance between raw outputs to recover the structure that predicts correlated failure.

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Axel Ahlqvist and colleagues at the UK AI Security Institute, Meridian and Anthropic attack evaluation awareness, the problem that capable models can tell when they are being tested rather than deployed, which weakens any conclusion a safety evaluation supports.

Evaluating and Improving LLM Self-Modeling
Siqi Zeng, Andre N. Assis and Rowan Wang, working through the Anthropic Fellows Program, measure whether a model can answer verifiable questions about its own behavior, and then try to train the ability in.

Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating
Heejin Choi at Yonsei University shows that the ranking of continual knowledge-updating methods reverses depending on when you evaluate and how much adapter capacity the baseline gets.