🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
614 papers · ReasoningClear filters →
Self-Discover

Self-Discover

Google's Self-Discover lets LLMs compose their own task-specific reasoning strategies from a small library of atomic reasoning modules, at dramatically lower inference cost than self-consistency.

505Reasoning
DeepSeekMath

DeepSeekMath

DeepSeek releases DeepSeekMath 7B, a math-specialized LLM that closes much of the gap to GPT-4 and Gemini-Ultra on MATH by combining better data and a new RL objective.

506Reasoning
LLMs for Table Processing: A Survey

LLMs for Table Processing: A Survey

A survey covering how LLMs and VLMs are used across the full spectrum of table-processing tasks, from classic TableQA to spreadsheet manipulation.

507Evaluation
LLMs for Mathematical Reasoning

LLMs for Mathematical Reasoning

A survey of the fast-growing literature on using LLMs for mathematical reasoning, from arithmetic word problems to theorem proving.

508Reasoning
Compression Algorithms for LLMs

Compression Algorithms for LLMs

A survey covering the main families of LLM compression techniques and when each one is appropriate.

509Efficiency
Knowledge Fusion of LLMs (FuseLLM)

Knowledge Fusion of LLMs (FuseLLM)

FuseLLM proposes fusing the capabilities of multiple existing LLMs into a single target model by distilling their output distributions rather than retraining from scratch.

510Training
AlphaGeometry

AlphaGeometry

DeepMind's AlphaGeometry is a theorem prover that solves Olympiad-level geometry problems at near gold-medallist performance, and crucially, without needing any human demonstrations.

511Reasoning
ReFT (Reinforced Fine-Tuning)

ReFT (Reinforced Fine-Tuning)

ByteDance's ReFT enhances LLM reasoning by combining supervised fine-tuning with online RL that samples alternative reasoning paths, without a learned reward model.

512Reasoning
Overview of LLMs for Evaluation

Overview of LLMs for Evaluation

A thorough survey of LLM-as-a-Judge and LLM-based evaluation methodologies, mapping strengths, limitations, and open problems.

513Evaluation
Patchscopes

Patchscopes

Patchscopes is a general framework for inspecting and intervening on LLM internals by "patching" hidden representations into a second inference pass.

514Reasoning
Easy-to-Hard Generalization

Easy-to-Hard Generalization

UNC researchers show that LLMs often generalize well from easy training data to hard evaluation data, with implications for scalable oversight.

515Evaluation
Chain-of-Table

Chain-of-Table

Google's Chain-of-Table prompts LLMs to iteratively transform a complex table step-by-step to answer questions reliably, extending CoT reasoning to tabular data.

516Reasoning
Mitigating Hallucination in LLMs

Mitigating Hallucination in LLMs

A survey cataloging 32 hallucination-mitigation techniques and organizing them into a practical taxonomy.

517Safety
LLM Augmented LLMs (CALM)

LLM Augmented LLMs (CALM)

Google's CALM composes a large anchor LLM with smaller specialist models via learned cross-attention, unlocking new capabilities without retraining either model.

518Code
SeeAct (GPT-4V as Generalist Web Agent)

SeeAct (GPT-4V as Generalist Web Agent)

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

519Agents
How Code Empowers LLMs

How Code Empowers LLMs

A survey on why training LLMs with code data produces capabilities well beyond coding itself.

520Agents
From Gemini to Q-Star

From Gemini to Q-Star

A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

521Multimodal
Fact Recalling in LLMs

Fact Recalling in LLMs

A mechanistic-interpretability study showing that early MLP layers function as a lookup table for factual recall.

522Safety
Generative AI for Math (OpenWebMath / MathPile)

Generative AI for Math (OpenWebMath / MathPile)

Releases a diverse, high-quality math-centric corpus of ~9.5B tokens designed for training math-capable foundation models.

523Reasoning
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

524Reasoning
Gemini vs GPT-4V

Gemini vs GPT-4V

A qualitative side-by-side comparison of Gemini and GPT-4V across vision-language tasks, documenting systematic behavioral differences.

525Multimodal
ReST Meets ReAct

ReST Meets ReAct

Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

526Agents
Mathematical LLMs Survey

Mathematical LLMs Survey

A survey on the progress of LLMs on mathematical reasoning tasks, covering methods, benchmarks, and open problems.

527Reasoning
Beyond Human Data (ReST-EM)

Beyond Human Data (ReST-EM)

DeepMind's ReST-EM shows that model-generated data plus a reward function can substantially reduce dependence on human-generated data.

528Reasoning
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026