🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization

Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization

Dong-Jae Koh, Young-Kyoon Suh and colleagues (Kyungpook National University) introduce DNBENCH for database normalization from 1NF to BCNF and a multi-agent method, MARS, that improves the benchmark score by 82.0% over a single prompt.

433Agents
MAPLE: Memory-Augmented Planning with Language and Evolution

MAPLE: Memory-Augmented Planning with Language and Evolution

Kesheng Chen, Yamin Hu and Wenjian Luo (Harbin Institute of Technology, Shenzhen) build MAPLE, an optimization agent that keeps an executable model of the problem and updates it across successive natural-language change requests.

434Agents
How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

Jeongyeon Kim and John Mitchell (Stanford University) measure how well multi-agent LLM pipelines perform qualitative coding, where agents code independently, debate and reconcile, and identify the dataset and process factors that decide accuracy.

435Agents
MindTopo: Can Foundation Models Reason in Topological Space?

MindTopo: Can Foundation Models Reason in Topological Space?

Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

436Reasoning
Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

Dong Li, Biqing Qi and colleagues (Shanghai AI Laboratory with Harbin Institute of Technology and others) build ARCHE, an agent that proposes reaction mechanisms, runs computational chemistry workflows to test them and revises conclusions from the computed evidence.

437Agents
When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

Ken Chen, Saman Halgamuge and colleagues (University of Melbourne) resolve disagreement between LLM agents by comparing each agent's forward answer with a posterior computed by Bayesian backward reasoning, which is less likely to share the same errors.

438Agents
Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

Shubham Agarwal, Alexander Krentsel, Shu Liu, Mert Cemri and colleagues (UC Berkeley with Google and UC Santa Cruz, including Ion Stoica, Matei Zaharia and Sylvia Ratnasamy) build Inductive Deductive Synthesis, an agentic LLM system that writes a distributed system's implementation and its mechanized correctness proof together, and it completes all seven key-value-store specifications it was given.

439Agents
The Time is Here for Just-in-Time Systems: Challenges and Opportunities

The Time is Here for Just-in-Time Systems: Challenges and Opportunities

Shu Liu, Alexander Krentsel, Shubham Agarwal, Mert Cemri and colleagues (UC Berkeley with Bespoke Labs) argue that coding agents make it practical to synthesize a core system from scratch for each deployment, and they present Jitskit, a pipeline that builds key-value stores specialized to one workload, one set of resource limits and one set of required guarantees.

440Agents
Artificial Id: Drive and Persistent Alignment in Agentic AI

Artificial Id: Drive and Persistent Alignment in Agentic AI

Yakov Pyotr Shkolnikov (independent researcher) proposes an artificial id, an internal adaptive drive that decides whether an agent's behavior should continue, stop or change, and argues that alignment must then apply to the continuing system rather than single responses.

441Safety
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

This survey splits recursive self-improvement into stages of autonomy, from executing improvements someone else designed up to improving the improvement process itself, which gives a concrete way to check what a claimed self-improving agent actually automates. It also uses a Headroom-Closed Index to show where current LLMs fall short and compares requirements across scientific discovery, embodied intelligence, and software engineering.

442Agents
terms.txt: A Consent and Compensation Protocol for Agentic Web Access

terms.txt: A Consent and Compensation Protocol for Agentic Web Access

Rajarshi Chowdhury (independent researcher) specifies terms.txt, a robots.txt-style file that states per-path, per-purpose access terms for AI crawlers and agents, together with a signed negotiation and receipt exchange that the origin server enforces.

443Agents
Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

Susheel Suresh and colleagues at Microsoft give the memory-curator agent in a GitHub Copilot harness read-only tools to check candidate memories against the live environment before they are saved, which roughly doubles pass rate on a database exploration benchmark and halves cost.

444Agents
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

Pingchen Lu, Zhongxiang Dai and colleagues (CUHK-Shenzhen, Tianjin University and NUS) treat agent skill optimization as a budgeted sequential decision problem and use a contextual bandit to decide which candidate skills are worth evaluating.

445Agents
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Minghao Guo, Meng Cao and colleagues (MBZUAI and USTC) build Mr.LHDR, a deep research benchmark whose questions require long chains of dependent, multimodal evidence, and find that the best system fully completes only about a third of them.

446Multimodal
But How Would AI Agents Run a Town's Economy?

But How Would AI Agents Run a Town's Economy?

Sajal Regmi, Siddhartha Pudasaini and Chetan Phakami Pun (Karela Technologies) put 100 memory-equipped LLM agents in charge of a closed town economy for up to 26 simulated weeks and find that money stops circulating: demand shocks raise revenue but wages and prices barely move.

447Agents
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Junyao Yang, Yucheng Shi and colleagues at Tencent's Hy Foundation Model Frontier team (with NUS, Georgia, Indiana and Maryland) train T1, a 122B-total MoE model, with RL in a real cloud shell for up to 300+ tool-call turns per task, and raise Terminal-Bench 2.1 from 43.8% to 64.0%.

448Reinforcement Learning
What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

Haseeb Mohammed Afsar (independent researcher) probes a seeded random sample of 400 servers from the 24,135-server MCP registry and finds that fewer than half start, and that popular tool-use benchmarks contain far more duplicated tools than real MCP servers do.

449Evaluation
ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI

ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI

Zhengran Ji, Jonathan Hyun and Boyuan Chen (Duke University) apply human organization theory to build task-specific hierarchies for teams of up to 50 embodied LLM agents, and beat four prior multi-agent frameworks across 25 wildfire-response missions.

450Robotics
Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender System

Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender System

Ming Li, Dai Li and colleagues at Meta build Auto-RecSys, an autonomous research agent system that runs multi-day experiments on industry-scale recommendation models, built around three harness designs and two self-improvement loops.

451Agents
Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

Ruiqing Yue, Yu Cui and colleagues (Chinese Academy of Sciences and Beijing Institute of Technology) speed up self-evolving agent harnesses by separating failures caused by the model from failures caused by the harness, and repairing only the recurring harness-level ones.

452Agents
Memory Compression for High-Fanout Agent Sandboxes

Memory Compression for High-Fanout Agent Sandboxes

Mengming Li, Ceyu Xu and colleagues (HKUST) build AgentZip, a memory compression system for agent workloads that spawn many concurrent sandboxes from a shared template, and cut sandbox memory by up to 8.7x.

453Memory
BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

Shenghan Zheng and Christophe Hauser (Dartmouth College), with Dawn Song (UC Berkeley) and collaborators at Amazon, BenchFlow and several universities, build BenchShield, an instrumentation layer that detects reward hacking in agent benchmarks from a formal model of each evaluation's reward-relevant events.

454Evaluation
Kernel-Managed Shared Memory for System-Wide Personalization

Kernel-Managed Shared Memory for System-Wide Personalization

Ryan Lum and Yongfeng Zhang (Rutgers University) move memory retrieval, privacy enforcement and prompt injection out of individual agents and into the agent-system kernel, and evaluate the design on AIOS across three assistant models and 1,800 trials.

455Memory
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

Xing Zhang, Guanghui Wang, Yanwei Cui, Mengdie Flora Wang and Peiyang He (AWS Generative AI Innovation Center) replace the generative manager in a compound LLM system with a defined operator, and show it beats generative managers on three held-out benchmarks.

456Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026