AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
ChatDev (Communicative Agents for Software Development)
ChatDev is a virtual chat-powered software company where LLM agents take on roles in a waterfall-model dev process.

OPRO (LLMs as Optimizers)
DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

Cognitive Architectures for Language Agents (CoALA)
Princeton proposes CoALA, a systematic framework for understanding and building language agents.

LLM-Based Autonomous Agents Survey
A comprehensive survey of LLM-based autonomous agents covering construction and applications.

Prompt2Model
CMU's Prompt2Model automates the path from a natural-language task description to a deployable small special-purpose model.

D-Bot (LLMs as Database Administrators)
Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

AgentBench
Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

LLMs for HVAC Control
Microsoft applies LLMs to industrial control tasks (HVAC for buildings), comparing against RL baselines.

ToolLLM
Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

MetaGPT
MetaGPT is a multi-agent framework that encodes standardized operating procedures (SOPs) for complex problem solving.

Dynalang (Agents Model the World with Language)
UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

WavJourney
Leverages LLMs to orchestrate audio generation models for compositional storytelling.

Generative TV & Showrunner Agents
Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

RoboCat
DeepMind's self-improving foundation agent that operates different robotic arms from as few as 100 demonstrations.

Mind2Web
A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

Thought Cloning
Imitation learning framework that learns to think as well as act.

Voyager
An LLM-powered embodied lifelong learning agent in Minecraft exploring autonomously.

Gorilla
A fine-tuned LLaMA-based model that surpasses GPT-4 on API call generation.

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
A framework inferring tool sequences for compositional reasoning.

Generative Agents: Interactive Simulacra of Human Behavior
Stanford/Google's landmark paper on LLM-powered social simulations.

Emergent Autonomous Scientific Research Capabilities of LLMs
An agent combining LLMs for autonomous scientific experiments.

ChemCrow: Augmenting LLMs with Chemistry Tools
An LLM chemistry agent with 13 expert-designed tools.

OpenAGI: When LLM Meets Domain Experts
An open-source research platform for LLM agents manipulating domain expert models.