🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
520 papers · 2024Clear filters →
TestGen-LLM

TestGen-LLM

Meta's TestGen-LLM uses LLMs to improve existing human-written tests - augmenting coverage rather than generating tests from scratch - while rigorously filtering LLM output for quality.

457Safety
ChemLLM

ChemLLM

ChemLLM is a chemistry-specialized LLM with a matched dataset (ChemData) and benchmark (ChemBench) for evaluating chemistry-specific capability.

458Evaluation
Survey of LLMs

Survey of LLMs

A survey that maps the landscape of the three dominant LLM families - GPT, Llama, and PaLM - and the shared toolbox used to build and augment them.

459Evaluation
LLM Agents Can Autonomously Hack Websites

LLM Agents Can Autonomously Hack Websites

The paper shows GPT-4 agents with tool use and long context can autonomously exploit real websites, including performing blind SQL injection and schema extraction.

460Agents
Grandmaster-Level Chess Without Search

Grandmaster-Level Chess Without Search

DeepMind shows that a 270M-parameter transformer trained purely with supervised learning on Stockfish-generated data reaches grandmaster-level chess without any search at inference time.

461Data
AnyTool

AnyTool

AnyTool is a training-free LLM agent that scales tool-use to 16K+ Rapid APIs through a hierarchical retriever and a self-reflective solver.

462Agents
Phase Transition in Dot-Product Attention

Phase Transition in Dot-Product Attention

A theoretical paper that analyzes a solvable low-rank tied-QK attention model and uncovers a data-driven phase transition between positional and semantic attention regimes.

463Architecture
Indirect Reasoning with LLMs (DIR)

Indirect Reasoning with LLMs (DIR)

Direct-Indirect Reasoning augments standard CoT with contrapositive and proof-by-contradiction templates, giving LLMs an explicit way to attack problems they can't solve forward.

464Reasoning
ALOHA 2

ALOHA 2

ALOHA 2 is a refreshed low-cost bimanual teleoperation platform from Stanford/DeepMind, designed for large-scale robot-learning data collection.

465Robotics
More Agents Is All You Need

More Agents Is All You Need

The paper shows that simply running more independent LLM agents and voting produces reliable scaling gains across tasks, without any method changes.

466Agents
Self-Discover

Self-Discover

Google's Self-Discover lets LLMs compose their own task-specific reasoning strategies from a small library of atomic reasoning modules, at dramatically lower inference cost than self-consistency.

467Reasoning
DeepSeekMath

DeepSeekMath

DeepSeek releases DeepSeekMath 7B, a math-specialized LLM that closes much of the gap to GPT-4 and Gemini-Ultra on MATH by combining better data and a new RL objective.

468Reasoning
LLMs for Table Processing: A Survey

LLMs for Table Processing: A Survey

A survey covering how LLMs and VLMs are used across the full spectrum of table-processing tasks, from classic TableQA to spreadsheet manipulation.

469Evaluation
LLM-based Multi-Agent Systems Survey

LLM-based Multi-Agent Systems Survey

A survey of the fast-growing LLM-based multi-agent systems space, covering both problem-solving applications and "world simulation" research.

470Agents
OLMo

OLMo

Allen AI releases OLMo, a truly open 7B-parameter LLM shipped with training code, pretraining data, full weights, evaluation tooling, and fine-tuning recipes - an answer to the "open-weights but closed-pipeline" releases dominating the space.

471Training
Advances in Multimodal LLMs

Advances in Multimodal LLMs

A comprehensive survey mapping design choices for architecture and training pipeline around multimodal large language models (MLLMs).

472Multimodal
Corrective RAG (CRAG)

Corrective RAG (CRAG)

CRAG adds a self-correcting loop around retrieval so a RAG system can detect and repair bad retrievals instead of feeding them straight into generation.

473Retrieval
LLMs for Mathematical Reasoning

LLMs for Mathematical Reasoning

A survey of the fast-growing literature on using LLMs for mathematical reasoning, from arithmetic word problems to theorem proving.

474Reasoning
Compression Algorithms for LLMs

Compression Algorithms for LLMs

A survey covering the main families of LLM compression techniques and when each one is appropriate.

475Efficiency
MoE-LLaVA

MoE-LLaVA

MoE-LLaVA applies Mixture-of-Experts tuning to the LLaVA vision-language architecture, getting a sparse model with dramatically fewer active parameters at the same compute cost.

476Architecture
Rephrasing the Web (WRAP)

Rephrasing the Web (WRAP)

WRAP uses an off-the-shelf instruction-tuned model to paraphrase web documents into styles like "Wikipedia" or "question-answer format" and trains on the mixture of real + synthetic rephrases.

477Training
The Power of Noise: Redefining Retrieval in RAG

The Power of Noise: Redefining Retrieval in RAG

A study stress-testing the retriever component of RAG systems with surprising results about what actually helps generation.

478Retrieval
Hallucination in LVLMs

Hallucination in LVLMs

A survey specifically scoped to hallucination in Large Vision-Language Models, a phenomenon that differs substantially from text-only LLM hallucination.

479Safety
SliceGPT

SliceGPT

Microsoft's SliceGPT is a post-training LLM compression technique that literally slices rows and columns out of weight matrices while preserving zero-shot quality.

480Efficiency
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026