🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
623 papers · AgentsClear filters →
The Dawn of LMMs (GPT-4V Deep Dive)

The Dawn of LMMs (GPT-4V Deep Dive)

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.

577Multimodal
Self-Taught Optimizer (STOP)

Self-Taught Optimizer (STOP)

Proposes recursively self-improving code generation where an LLM-scaffolded program improves itself.

578Code
Qwen

Qwen

Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

579Training
Compositional Foundation Models (HiP)

Compositional Foundation Models (HiP)

Proposes foundation models that compose multiple expert foundation models trained on different modalities to solve long-horizon goals.

580Architecture
OWL (LLMs for IT Operations)

OWL (LLMs for IT Operations)

Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

581Evaluation
The Rise and Potential of LLM-Based Agents

The Rise and Potential of LLM-Based Agents

A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

582Agents
Agents Library

Agents Library

An open-source library for building autonomous language agents with first-class support for planning, memory, tools, and multi-agent communication.

583Agents
ChatDev (Communicative Agents for Software Development)

ChatDev (Communicative Agents for Software Development)

ChatDev is a virtual chat-powered software company where LLM agents take on roles in a waterfall-model dev process.

584Agents
GPT Solves Math Problems Without a Calculator

GPT Solves Math Problems Without a Calculator

Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

585Reasoning
OPRO (LLMs as Optimizers)

OPRO (LLMs as Optimizers)

DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

586Agents
Cognitive Architectures for Language Agents (CoALA)

Cognitive Architectures for Language Agents (CoALA)

Princeton proposes CoALA, a systematic framework for understanding and building language agents.

587Agents
Q-Transformer

Q-Transformer

Google's Q-Transformer is a scalable RL method for training multi-task robotic policies from large offline datasets.

588Robotics
LLM-Based Autonomous Agents Survey

LLM-Based Autonomous Agents Survey

A comprehensive survey of LLM-based autonomous agents covering construction and applications.

589Agents
Prompt2Model

Prompt2Model

CMU's Prompt2Model automates the path from a natural-language task description to a deployable small special-purpose model.

590Agents
Outlines (Efficient Guided Generation)

Outlines (Efficient Guided Generation)

A library for guided LLM text generation that enforces structural constraints with minimal overhead.

591Efficiency
D-Bot (LLMs as Database Administrators)

D-Bot (LLMs as Database Administrators)

Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

592Agents
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

593Agents
LLMs for HVAC Control

LLMs for HVAC Control

Microsoft applies LLMs to industrial control tasks (HVAC for buildings), comparing against RL baselines.

594Agents
ToolLLM

ToolLLM

Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

595Agents
MetaGPT

MetaGPT

MetaGPT is a multi-agent framework that encodes standardized operating procedures (SOPs) for complex problem solving.

596Agents
Dynalang (Agents Model the World with Language)

Dynalang (Agents Model the World with Language)

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

597Agents
WavJourney

WavJourney

Leverages LLMs to orchestrate audio generation models for compositional storytelling.

598Multimodal
Generative TV & Showrunner Agents

Generative TV & Showrunner Agents

Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

599Agents
A Survey on Evaluation of LLMs

A Survey on Evaluation of LLMs

A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.

600Evaluation
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026