🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
Dobb-E

Dobb-E

NYU's Dobb-E is an affordable household-manipulation robot that learns new tasks with just 5 minutes of user demonstrations.

49Robotics
Translatotron 3

Translatotron 3

Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.

50Data
System 2 Attention (S2A)

System 2 Attention (S2A)

Meta's S2A uses the LLM's own reasoning to decide what context actually matters, regenerating a clean prompt before the final response step.

51Reasoning
Advancing Long-Context LLMs

Advancing Long-Context LLMs

A survey of methodologies for improving Transformer long-context capability across pretraining, fine-tuning, and inference stages.

52Memory
Parallel Speculative Sampling

Parallel Speculative Sampling

Amazon researchers propose a parallel variant of speculative sampling that achieves significant LLM inference speedups with minimal extra parameters.

53Efficiency
Mirasol3B

Mirasol3B

Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

54Multimodal
Teaching Small LMs to Reason

Teaching Small LMs to Reason

An approach that teaches smaller language models to explicitly select among reasoning techniques for each problem.

55Reasoning
GPQA

GPQA

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

56Reasoning
Hitchhiker's Guide From CoT to Agents

Hitchhiker's Guide From CoT to Agents

A survey mapping the conceptual evolution from chain-of-thought reasoning to modern language-agent frameworks.

57Agents
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

58Agents
MedAgents

MedAgents

A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

59Reasoning
TÜLU 2

TÜLU 2

Allen AI's TÜLU 2 is a suite of improved open instruction-tuned LLMs and an accompanying study of adaptation best practices.

60Training
Emu Video and Emu Edit

Emu Video and Emu Edit

Meta releases Emu Video and Emu Edit, a pair of diffusion models targeting controlled text-to-video generation and instruction-based image editing.

61Multimodal
Chain-of-Note (CoN)

Chain-of-Note (CoN)

Tencent's Chain-of-Note adds an explicit note-taking step to RAG so the model can evaluate retrieved evidence before answering.

62Retrieval
LLMs for Scientific Discovery

LLMs for Scientific Discovery

A broad evaluation of GPT-4 across scientific disciplines including drug discovery, biology, and computational chemistry.

63Evaluation
Fine-Tuning LLMs for Factuality

Fine-Tuning LLMs for Factuality

Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

64Training
Contrastive Chain-of-Thought

Contrastive Chain-of-Thought

Proposes contrastive CoT prompting where models see both valid *and* invalid reasoning demonstrations to reduce reasoning errors.

65Reasoning
Survey on Language Models for Code

Survey on Language Models for Code

A comprehensive survey of LLMs for code covering 50+ models, 30+ evaluation tasks, and 500 related works.

66Evaluation
JARVIS-1

JARVIS-1

An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

67Agents
Learning to Filter Context for RAG (FILCO)

Learning to Filter Context for RAG (FILCO)

CMU's FILCO improves RAG by training a dedicated model to filter retrieved contexts before they reach the generator.

68Retrieval
MART (Multi-round Automatic Red-Teaming)

MART (Multi-round Automatic Red-Teaming)

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

69Safety
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

70Agents
Hallucination in LLMs Survey

Hallucination in LLMs Survey

A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

71Safety
Simplifying Transformer Blocks

Simplifying Transformer Blocks

Researchers show that many components of the standard transformer block can be removed with no loss in training speed or quality.

72Training
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026