🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
ChatDev (Communicative Agents for Software Development)

ChatDev (Communicative Agents for Software Development)

ChatDev is a virtual chat-powered software company where LLM agents take on roles in a waterfall-model dev process.

1105Agents
OPRO (LLMs as Optimizers)

OPRO (LLMs as Optimizers)

DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

1106Agents
Cognitive Architectures for Language Agents (CoALA)

Cognitive Architectures for Language Agents (CoALA)

Princeton proposes CoALA, a systematic framework for understanding and building language agents.

1107Agents
LLM-Based Autonomous Agents Survey

LLM-Based Autonomous Agents Survey

A comprehensive survey of LLM-based autonomous agents covering construction and applications.

1108Agents
Prompt2Model

Prompt2Model

CMU's Prompt2Model automates the path from a natural-language task description to a deployable small special-purpose model.

1109Agents
D-Bot (LLMs as Database Administrators)

D-Bot (LLMs as Database Administrators)

Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

1110Agents
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

1111Agents
LLMs for HVAC Control

LLMs for HVAC Control

Microsoft applies LLMs to industrial control tasks (HVAC for buildings), comparing against RL baselines.

1112Agents
ToolLLM

ToolLLM

Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

1113Agents
MetaGPT

MetaGPT

MetaGPT is a multi-agent framework that encodes standardized operating procedures (SOPs) for complex problem solving.

1114Agents
Dynalang (Agents Model the World with Language)

Dynalang (Agents Model the World with Language)

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

1115Agents
WavJourney

WavJourney

Leverages LLMs to orchestrate audio generation models for compositional storytelling.

1116Multimodal
Generative TV & Showrunner Agents

Generative TV & Showrunner Agents

Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

1117Agents
RoboCat

RoboCat

DeepMind's self-improving foundation agent that operates different robotic arms from as few as 100 demonstrations.

1118Robotics
Mind2Web

Mind2Web

A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

1119Agents
Thought Cloning

Thought Cloning

Imitation learning framework that learns to think as well as act.

1120Agents
Voyager

Voyager

An LLM-powered embodied lifelong learning agent in Minecraft exploring autonomously.

1121Agents
Gorilla

Gorilla

A fine-tuned LLaMA-based model that surpasses GPT-4 on API call generation.

1122Agents
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

A framework inferring tool sequences for compositional reasoning.

1123Reasoning
Generative Agents: Interactive Simulacra of Human Behavior

Generative Agents: Interactive Simulacra of Human Behavior

Stanford/Google's landmark paper on LLM-powered social simulations.

1124Agents
Emergent Autonomous Scientific Research Capabilities of LLMs

Emergent Autonomous Scientific Research Capabilities of LLMs

An agent combining LLMs for autonomous scientific experiments.

1125Agents
ChemCrow: Augmenting LLMs with Chemistry Tools

ChemCrow: Augmenting LLMs with Chemistry Tools

An LLM chemistry agent with 13 expert-designed tools.

1126Agents
OpenAGI: When LLM Meets Domain Experts

OpenAGI: When LLM Meets Domain Experts

An open-source research platform for LLM agents manipulating domain expert models.

1127Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026