AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

Robots That Ask for Help
A framework for calibrating LLM-based robot planners so they ask for help when uncertain.

InterCode
A framework treating interactive coding as a reinforcement learning environment.

Understanding Theory-of-Mind in LLMs with LLMs
A framework for procedurally generating ToM evaluations using LLMs themselves.

RoboCat
DeepMind's self-improving foundation agent that operates different robotic arms from as few as 100 demonstrations.

Unifying LLMs & Knowledge Graphs
A roadmap for combining LLMs with knowledge graphs for stronger reasoning.

Mind2Web
A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

AlphaDev
DeepMind's deep RL agent discovering faster sorting algorithms from scratch, now in LLVM.

Augmenting LLMs with Databases (ChatDB)
Combines an LLM with SQL databases as a symbolic memory framework.

Thought Cloning
Imitation learning framework that learns to think as well as act.

Voyager
An LLM-powered embodied lifelong learning agent in Minecraft exploring autonomously.

Gorilla
A fine-tuned LLaMA-based model that surpasses GPT-4 on API call generation.

StructGPT
A general framework for LLM reasoning over structured data.

TidyBot
Combines LLM-based planning and perception with few-shot summarization to infer user preferences.

Track Anything
An interactive tool for video object tracking and segmentation built on Segment Anything.

AudioGPT
Connects ChatGPT with audio foundational models for speech, music, sound, and talking head tasks.

Learning to Compress Prompts with Gist Tokens
Trains LMs to compress prompts into reusable "gist" tokens.

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
A framework inferring tool sequences for compositional reasoning.

Generative Agents: Interactive Simulacra of Human Behavior
Stanford/Google's landmark paper on LLM-powered social simulations.

Emergent Autonomous Scientific Research Capabilities of LLMs
An agent combining LLMs for autonomous scientific experiments.

ChemCrow: Augmenting LLMs with Chemistry Tools
An LLM chemistry agent with 13 expert-designed tools.

OpenAGI: When LLM Meets Domain Experts
An open-source research platform for LLM agents manipulating domain expert models.

Teaching Large Language Models to Self-Debug
Teaches LLMs to debug their own code via few-shot demonstrations.

MACHIAVELLI Benchmark
A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.