🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
614 papers · ReasoningClear filters →
D-Bot (LLMs as Database Administrators)

D-Bot (LLMs as Database Administrators)

Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

577Agents
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

578Agents
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

579Safety
Skeleton-of-Thought (SoT)

Skeleton-of-Thought (SoT)

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

580Reasoning
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

581Safety
RT-2

RT-2

Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.

582Robotics
FacTool

FacTool

A task- and domain-agnostic framework for factuality detection of LLM-generated text.

583Evaluation
How is ChatGPT's Behavior Changing Over Time?

How is ChatGPT's Behavior Changing Over Time?

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

584Evaluation
Measuring Faithfulness in Chain-of-Thought Reasoning

Measuring Faithfulness in Chain-of-Thought Reasoning

Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

585Reasoning
FLASK

FLASK

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.

586Evaluation
Claude 2

Claude 2

Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

587Safety
LLMs as General Pattern Machines

LLMs as General Pattern Machines

Demonstrates LLMs serve as general sequence modelers without additional training.

588Reasoning
Teaching Arithmetic to Small Transformers

Teaching Arithmetic to Small Transformers

Trains small transformers on chain-of-thought style data for arithmetic with large gains.

589Reasoning
LLMs as Effective Text Rankers

LLMs as Effective Text Rankers

A prompting technique that enables open-source LLMs to perform SOTA text ranking.

590Retrieval
Multimodal Generation with Frozen LLMs

Multimodal Generation with Frozen LLMs

Maps images to LLM token space enabling models like PaLM and GPT-4 to handle visual tasks without parameter updates.

591Multimodal
LeanDojo

LeanDojo

An open-source Lean playground consisting of toolkits, data, models, and benchmarks for theorem proving.

592Reasoning
Computer Vision Through the Lens of Natural Language

Computer Vision Through the Lens of Natural Language

A modular approach solving CV problems by routing through LLM reasoning.

593Multimodal
Understanding Theory-of-Mind in LLMs with LLMs

Understanding Theory-of-Mind in LLMs with LLMs

A framework for procedurally generating ToM evaluations using LLMs themselves.

594Evaluation
SequenceMatch

SequenceMatch

Formulates sequence generation as imitation learning, enabling backtracking via a backspace action.

595Reinforcement Learning
Unifying LLMs & Knowledge Graphs

Unifying LLMs & Knowledge Graphs

A roadmap for combining LLMs with knowledge graphs for stronger reasoning.

596Retrieval
Augmenting LLMs with Databases (ChatDB)

Augmenting LLMs with Databases (ChatDB)

Combines an LLM with SQL databases as a symbolic memory framework.

597Memory
Imitating Reasoning Process of Larger LLMs (Orca)

Imitating Reasoning Process of Larger LLMs (Orca)

Microsoft's 13B model that imitates GPT-4's reasoning traces.

598Reasoning
Let's Verify Step by Step

Let's Verify Step by Step

OpenAI's landmark paper on process reward models for mathematical reasoning.

599Reasoning
Thought Cloning

Thought Cloning

Imitation learning framework that learns to think as well as act.

600Agents
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026