AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization
Dong-Jae Koh, Young-Kyoon Suh and colleagues (Kyungpook National University) introduce DNBENCH for database normalization from 1NF to BCNF and a multi-agent method, MARS, that improves the benchmark score by 82.0% over a single prompt.

MAPLE: Memory-Augmented Planning with Language and Evolution
Kesheng Chen, Yamin Hu and Wenjian Luo (Harbin Institute of Technology, Shenzhen) build MAPLE, an optimization agent that keeps an executable model of the problem and updates it across successive natural-language change requests.

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding
Jeongyeon Kim and John Mitchell (Stanford University) measure how well multi-agent LLM pipelines perform qualitative coding, where agents code independently, debate and reconcile, and identify the dataset and process factors that decide accuracy.

MindTopo: Can Foundation Models Reason in Topological Space?
Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation
Dong Li, Biqing Qi and colleagues (Shanghai AI Laboratory with Harbin Institute of Technology and others) build ARCHE, an agent that proposes reaction mechanisms, runs computational chemistry workflows to test them and revises conclusions from the computed evidence.

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making
Ken Chen, Saman Halgamuge and colleagues (University of Melbourne) resolve disagreement between LLM agents by comparing each agent's forward answer with a posterior computed by Bayesian backward reasoning, which is less likely to share the same errors.

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
Shubham Agarwal, Alexander Krentsel, Shu Liu, Mert Cemri and colleagues (UC Berkeley with Google and UC Santa Cruz, including Ion Stoica, Matei Zaharia and Sylvia Ratnasamy) build Inductive Deductive Synthesis, an agentic LLM system that writes a distributed system's implementation and its mechanized correctness proof together, and it completes all seven key-value-store specifications it was given.

The Time is Here for Just-in-Time Systems: Challenges and Opportunities
Shu Liu, Alexander Krentsel, Shubham Agarwal, Mert Cemri and colleagues (UC Berkeley with Bespoke Labs) argue that coding agents make it practical to synthesize a core system from scratch for each deployment, and they present Jitskit, a pipeline that builds key-value stores specialized to one workload, one set of resource limits and one set of required guarantees.

Artificial Id: Drive and Persistent Alignment in Agentic AI
Yakov Pyotr Shkolnikov (independent researcher) proposes an artificial id, an internal adaptive drive that decides whether an agent's behavior should continue, stop or change, and argues that alignment must then apply to the continuing system rather than single responses.

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
This survey splits recursive self-improvement into stages of autonomy, from executing improvements someone else designed up to improving the improvement process itself, which gives a concrete way to check what a claimed self-improving agent actually automates. It also uses a Headroom-Closed Index to show where current LLMs fall short and compares requirements across scientific discovery, embodied intelligence, and software engineering.

terms.txt: A Consent and Compensation Protocol for Agentic Web Access
Rajarshi Chowdhury (independent researcher) specifies terms.txt, a robots.txt-style file that states per-path, per-purpose access terms for AI crawlers and agents, together with a signed negotiation and receipt exchange that the origin server enforces.

Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents
Susheel Suresh and colleagues at Microsoft give the memory-curator agent in a GitHub Copilot harness read-only tools to check candidate memories against the live environment before they are saved, which roughly doubles pass rate on a database exploration benchmark and halves cost.

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Pingchen Lu, Zhongxiang Dai and colleagues (CUHK-Shenzhen, Tianjin University and NUS) treat agent skill optimization as a budgeted sequential decision problem and use a contextual bandit to decide which candidate skills are worth evaluating.

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
Minghao Guo, Meng Cao and colleagues (MBZUAI and USTC) build Mr.LHDR, a deep research benchmark whose questions require long chains of dependent, multimodal evidence, and find that the best system fully completes only about a third of them.

But How Would AI Agents Run a Town's Economy?
Sajal Regmi, Siddhartha Pudasaini and Chetan Phakami Pun (Karela Technologies) put 100 memory-equipped LLM agents in charge of a closed town economy for up to 26 simulated weeks and find that money stops circulating: demand shocks raise revenue but wages and prices barely move.

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Junyao Yang, Yucheng Shi and colleagues at Tencent's Hy Foundation Model Frontier team (with NUS, Georgia, Indiana and Maryland) train T1, a 122B-total MoE model, with RL in a real cloud shell for up to 300+ tool-call turns per task, and raise Terminal-Bench 2.1 from 43.8% to 64.0%.

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead
Haseeb Mohammed Afsar (independent researcher) probes a seeded random sample of 400 servers from the 24,135-server MCP registry and finds that fewer than half start, and that popular tool-use benchmarks contain far more duplicated tools than real MCP servers do.

ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI
Zhengran Ji, Jonathan Hyun and Boyuan Chen (Duke University) apply human organization theory to build task-specific hierarchies for teams of up to 50 embodied LLM agents, and beat four prior multi-agent frameworks across 25 wildfire-response missions.

Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender System
Ming Li, Dai Li and colleagues at Meta build Auto-RecSys, an autonomous research agent system that runs multi-day experiments on industry-scale recommendation models, built around three harness designs and two self-improvement loops.

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents
Ruiqing Yue, Yu Cui and colleagues (Chinese Academy of Sciences and Beijing Institute of Technology) speed up self-evolving agent harnesses by separating failures caused by the model from failures caused by the harness, and repairing only the recurring harness-level ones.

Memory Compression for High-Fanout Agent Sandboxes
Mengming Li, Ceyu Xu and colleagues (HKUST) build AgentZip, a memory compression system for agent workloads that spawn many concurrent sandboxes from a shared template, and cut sandbox memory by up to 8.7x.

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
Shenghan Zheng and Christophe Hauser (Dartmouth College), with Dawn Song (UC Berkeley) and collaborators at Amazon, BenchFlow and several universities, build BenchShield, an instrumentation layer that detects reward hacking in agent benchmarks from a formal model of each evaluation's reward-relevant events.

Kernel-Managed Shared Memory for System-Wide Personalization
Ryan Lum and Yongfeng Zhang (Rutgers University) move memory retrieval, privacy enforcement and prompt injection out of individual agents and into the agent-system kernel, and evaluate the design on AIOS across three assistant models and 1,800 trials.

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
Xing Zhang, Guanghui Wang, Yanwei Cui, Mengdie Flora Wang and Peiyang He (AWS Generative AI Innovation Center) replace the generative manager in a compound LLM system with a defined operator, and show it beats generative managers on three held-out benchmarks.