🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
614 papers · ReasoningClear filters →
LLM Self-Explanations

LLM Self-Explanations

Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

553Safety
Hypothesis Search (LLMs Can Learn Rules)

Hypothesis Search (LLMs Can Learn Rules)

A two-stage framework where the LLM learns a rule library for reasoning.

554Reasoning
Meta Chain-of-Thought Prompting (Meta-CoT)

Meta Chain-of-Thought Prompting (Meta-CoT)

A generalizable CoT framework that selects domain-appropriate reasoning patterns for the task at hand.

555Reasoning
MemWalker

MemWalker

MemWalker treats the LLM as an interactive agent that traverses a tree-structured summary of long text.

556Memory
LLMs Represent Space and Time

LLMs Represent Space and Time

MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

557Reasoning
The Dawn of LMMs (GPT-4V Deep Dive)

The Dawn of LMMs (GPT-4V Deep Dive)

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.

558Multimodal
Training LLMs with Pause Tokens

Training LLMs with Pause Tokens

CMU shows that adding a learnable `<pause>` token during both pretraining and fine-tuning gives the model extra "thinking time" and improves reasoning.

559Training
Analogical Prompting

Analogical Prompting

Google's Analogical Prompting guides LLM reasoning by having the model self-generate relevant exemplars on the fly.

560Reasoning
Effective Long-Context Scaling (Meta)

Effective Long-Context Scaling (Meta)

Meta proposes a 70B long-context LLM that surpasses GPT-3.5-turbo-16k on long-context benchmarks.

561Memory
Boolformer

Boolformer

The first Transformer trained to perform end-to-end symbolic regression of Boolean functions.

562Reasoning
Logical Chain-of-Thought (LogiCoT)

Logical Chain-of-Thought (LogiCoT)

A neurosymbolic framework that verifies and revises zero-shot CoT reasoning using symbolic-logic principles.

563Reasoning
Chain-of-Verification (CoVe)

Chain-of-Verification (CoVe)

Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

564Safety
Contrastive Decoding for Reasoning

Contrastive Decoding for Reasoning

Shows that contrastive decoding, a simple inference-time technique, substantially improves reasoning in large LLMs.

565Reasoning
Textbooks Are All You Need II (phi-1.5)

Textbooks Are All You Need II (phi-1.5)

Microsoft's phi-1.5 demonstrates that a 1.3B model trained on "textbook-quality" synthetic data rivals much larger models on reasoning.

566Data
MAmmoTH

MAmmoTH

An open-source LLM family specialized for general mathematical problem solving.

567Reasoning
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

568Reinforcement Learning
GPT Solves Math Problems Without a Calculator

GPT Solves Math Problems Without a Calculator

Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

569Reasoning
OPRO (LLMs as Optimizers)

OPRO (LLMs as Optimizers)

DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

570Agents
Cognitive Architectures for Language Agents (CoALA)

Cognitive Architectures for Language Agents (CoALA)

Princeton proposes CoALA, a systematic framework for understanding and building language agents.

571Agents
Graph of Thoughts (GoT)

Graph of Thoughts (GoT)

Generalizes Chain-of-Thought and Tree-of-Thought by modeling LLM reasoning as an arbitrary graph.

572Reasoning
FacTool

FacTool

A tool-augmented framework for detecting factual errors in LLM-generated text.

573Retrieval
LegalBench

LegalBench

A collaboratively constructed benchmark for measuring legal reasoning in LLMs.

574Evaluation
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

575Efficiency
GPT-4 Code Interpreter for Math

GPT-4 Code Interpreter for Math

A zero-shot prompting technique for GPT-4 Code Interpreter that dramatically boosts math-reasoning accuracy via code self-verification.

576Reasoning
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026