🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
623 papers · AgentsClear filters →
Robots That Ask for Help

Robots That Ask for Help

A framework for calibrating LLM-based robot planners so they ask for help when uncertain.

601Robotics
InterCode

InterCode

A framework treating interactive coding as a reinforcement learning environment.

602Reinforcement Learning
Understanding Theory-of-Mind in LLMs with LLMs

Understanding Theory-of-Mind in LLMs with LLMs

A framework for procedurally generating ToM evaluations using LLMs themselves.

603Evaluation
RoboCat

RoboCat

DeepMind's self-improving foundation agent that operates different robotic arms from as few as 100 demonstrations.

604Robotics
Unifying LLMs & Knowledge Graphs

Unifying LLMs & Knowledge Graphs

A roadmap for combining LLMs with knowledge graphs for stronger reasoning.

605Retrieval
Mind2Web

Mind2Web

A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

606Agents
AlphaDev

AlphaDev

DeepMind's deep RL agent discovering faster sorting algorithms from scratch, now in LLVM.

607Reinforcement Learning
Augmenting LLMs with Databases (ChatDB)

Augmenting LLMs with Databases (ChatDB)

Combines an LLM with SQL databases as a symbolic memory framework.

608Memory
Thought Cloning

Thought Cloning

Imitation learning framework that learns to think as well as act.

609Agents
Voyager

Voyager

An LLM-powered embodied lifelong learning agent in Minecraft exploring autonomously.

610Agents
Gorilla

Gorilla

A fine-tuned LLaMA-based model that surpasses GPT-4 on API call generation.

611Agents
StructGPT

StructGPT

A general framework for LLM reasoning over structured data.

612Reasoning
TidyBot

TidyBot

Combines LLM-based planning and perception with few-shot summarization to infer user preferences.

613Robotics
Track Anything

Track Anything

An interactive tool for video object tracking and segmentation built on Segment Anything.

614Multimodal
AudioGPT

AudioGPT

Connects ChatGPT with audio foundational models for speech, music, sound, and talking head tasks.

615Multimodal
Learning to Compress Prompts with Gist Tokens

Learning to Compress Prompts with Gist Tokens

Trains LMs to compress prompts into reusable "gist" tokens.

616Efficiency
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

A framework inferring tool sequences for compositional reasoning.

617Reasoning
Generative Agents: Interactive Simulacra of Human Behavior

Generative Agents: Interactive Simulacra of Human Behavior

Stanford/Google's landmark paper on LLM-powered social simulations.

618Agents
Emergent Autonomous Scientific Research Capabilities of LLMs

Emergent Autonomous Scientific Research Capabilities of LLMs

An agent combining LLMs for autonomous scientific experiments.

619Agents
ChemCrow: Augmenting LLMs with Chemistry Tools

ChemCrow: Augmenting LLMs with Chemistry Tools

An LLM chemistry agent with 13 expert-designed tools.

620Agents
OpenAGI: When LLM Meets Domain Experts

OpenAGI: When LLM Meets Domain Experts

An open-source research platform for LLM agents manipulating domain expert models.

621Agents
Teaching Large Language Models to Self-Debug

Teaching Large Language Models to Self-Debug

Teaches LLMs to debug their own code via few-shot demonstrations.

622Code
MACHIAVELLI Benchmark

MACHIAVELLI Benchmark

A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.

623Evaluation
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026