🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,762
Papers
176
Weekly issues
2023
Since
862 papers · EvaluationClear filters →
PlanGPT

PlanGPT

PlanGPT is a domain-specialized LLM framework for urban and spatial planning, built in collaboration with the Chinese Academy of Urban Planning.

649Agents
Gemma

Gemma

Google DeepMind releases Gemma, a family of open models (2B and 7B) built from the same research stack as Gemini and shipped with both base and instruction-tuned variants.

650Reinforcement Learning
GRIT

GRIT

GRIT (Generative Representational Instruction Tuning) trains a single LLM to handle both generative and embedding tasks, switching behavior based on instructions.

651Retrieval
Back to Basics: Revisiting REINFORCE in RLHF

Back to Basics: Revisiting REINFORCE in RLHF

Cohere researchers argue that PPO is overkill for RLHF and that a simpler REINFORCE-style estimator works better in practice.

652Reinforcement Learning
Recurrent Memory Finds What LLMs Miss

Recurrent Memory Finds What LLMs Miss

Introduces BABILong, a new long-context benchmark, and shows that transformers with recurrent memory can handle sequences far beyond vanilla LLMs.

653Memory
Chain-of-Thought Reasoning Without Prompting

Chain-of-Thought Reasoning Without Prompting

DeepMind shows that LLMs often *already* emit chain-of-thought reasoning in alternative decoding paths, and that selecting those paths via confidence lifts reasoning accuracy with no prompt engineering.

654Reasoning
OpenCodeInterpreter

OpenCodeInterpreter

OpenCodeInterpreter is an open-source family of code-execution LLM systems that iteratively refine code using runtime feedback, closing the gap with GPT-4's proprietary Code Interpreter.

655Agents
Gemini 1.5

Gemini 1.5

Google DeepMind's Gemini 1.5 is a multimodal MoE LLM that scales context to 1M tokens (10M in research settings) while matching or surpassing Gemini 1.0 Ultra on standard benchmarks.

656Architecture
V-JEPA

V-JEPA

Meta's V-JEPA learns visual representations by predicting features in masked video regions, without pretrained image encoders, text, negatives, or reconstruction.

657Training
Large World Model (LWM)

Large World Model (LWM)

UC Berkeley's LWM is an open 7B multimodal model trained on long videos and books that handles context windows up to 1M tokens via RingAttention.

658Memory
OS-Copilot

OS-Copilot

OS-Copilot is a framework for building generalist computer agents that use full OS primitives (browser, terminal, files, multimedia, third-party apps) rather than just web DOMs.

659Agents
TestGen-LLM

TestGen-LLM

Meta's TestGen-LLM uses LLMs to improve existing human-written tests - augmenting coverage rather than generating tests from scratch - while rigorously filtering LLM output for quality.

660Safety
ChemLLM

ChemLLM

ChemLLM is a chemistry-specialized LLM with a matched dataset (ChemData) and benchmark (ChemBench) for evaluating chemistry-specific capability.

661Evaluation
Survey of LLMs

Survey of LLMs

A survey that maps the landscape of the three dominant LLM families - GPT, Llama, and PaLM - and the shared toolbox used to build and augment them.

662Evaluation
AnyTool

AnyTool

AnyTool is a training-free LLM agent that scales tool-use to 16K+ Rapid APIs through a hierarchical retriever and a self-reflective solver.

663Agents
Indirect Reasoning with LLMs (DIR)

Indirect Reasoning with LLMs (DIR)

Direct-Indirect Reasoning augments standard CoT with contrapositive and proof-by-contradiction templates, giving LLMs an explicit way to attack problems they can't solve forward.

664Reasoning
More Agents Is All You Need

More Agents Is All You Need

The paper shows that simply running more independent LLM agents and voting produces reliable scaling gains across tasks, without any method changes.

665Agents
LLMs for Table Processing: A Survey

LLMs for Table Processing: A Survey

A survey covering how LLMs and VLMs are used across the full spectrum of table-processing tasks, from classic TableQA to spreadsheet manipulation.

666Evaluation
LLM-based Multi-Agent Systems Survey

LLM-based Multi-Agent Systems Survey

A survey of the fast-growing LLM-based multi-agent systems space, covering both problem-solving applications and "world simulation" research.

667Agents
OLMo

OLMo

Allen AI releases OLMo, a truly open 7B-parameter LLM shipped with training code, pretraining data, full weights, evaluation tooling, and fine-tuning recipes - an answer to the "open-weights but closed-pipeline" releases dominating the space.

668Training
Advances in Multimodal LLMs

Advances in Multimodal LLMs

A comprehensive survey mapping design choices for architecture and training pipeline around multimodal large language models (MLLMs).

669Multimodal
Corrective RAG (CRAG)

Corrective RAG (CRAG)

CRAG adds a self-correcting loop around retrieval so a RAG system can detect and repair bad retrievals instead of feeding them straight into generation.

670Retrieval
LLMs for Mathematical Reasoning

LLMs for Mathematical Reasoning

A survey of the fast-growing literature on using LLMs for mathematical reasoning, from arithmetic word problems to theorem proving.

671Reasoning
MoE-LLaVA

MoE-LLaVA

MoE-LLaVA applies Mixture-of-Experts tuning to the LLaVA vision-language architecture, getting a sparse model with dramatically fewer active parameters at the same compute cost.

672Architecture
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026