🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
520 papers · 2024Clear filters →
DROID

DROID

DROID is an open-source robot manipulation dataset that dramatically expands the diversity of real-world robot demonstrations available for imitation-learning research.

409Robotics
RAFT: Retrieval-Augmented Fine-Tuning

RAFT: Retrieval-Augmented Fine-Tuning

RAFT is a fine-tuning recipe that teaches LLMs to handle distractor documents during RAG and to answer with CoT-style citations to retrieved passages.

410Retrieval
SIMA

SIMA

DeepMind's Scalable Instructable Multiworld Agent (SIMA) is a generalist AI agent that follows natural-language instructions across nine commercial 3D video games like No Man's Sky, Teardown, Valheim, and Space Engineers.

411Agents
Retrieval Augmented Thoughts (RAT)

Retrieval Augmented Thoughts (RAT)

RAT augments chain-of-thought by iteratively rewriting each reasoning step using retrieved context, sharply reducing hallucination on long-horizon generation tasks.

412Retrieval
Quiet-STaR

Quiet-STaR

Quiet-STaR generalizes the Self-Taught Reasoner (STaR) so that a language model learns to generate internal rationales between every token, not just for explicit QA problems.

413Reasoning
Knowledge Conflicts for LLMs

Knowledge Conflicts for LLMs

A survey that maps the landscape of knowledge conflicts in LLMs, covering how they arise, how models behave under them, and how to mitigate them.

414Memory
Stealing Part of a Production Language Model

Stealing Part of a Production Language Model

The paper demonstrates the first practical attack that extracts the embedding-projection layer of production LLMs through their ordinary logit APIs.

415Safety
Branch-Train-MiX (BTX)

Branch-Train-MiX (BTX)

Meta's BTX produces a single Mixture-of-Experts LLM by first training specialized experts in parallel and then mixing them, sidestepping the high cost of training one big generalist.

416Training
LLMs Predict Neuroscience Results (BrainBench)

LLMs Predict Neuroscience Results (BrainBench)

BrainBench asks both LLMs and human experts to predict the outcomes of neuroscience experiments from their abstracts, and finds LLMs outperform experts.

417Evaluation
C4AI Command-R

C4AI Command-R

Cohere for AI releases Command-R, a 35B open-weight LLM tuned specifically for retrieval-augmented generation, tool use, and multilingual workflows.

418Retrieval
Is Cosine-Similarity Really About Similarity?

Is Cosine-Similarity Really About Similarity?

This paper argues that cosine similarity between learned embeddings does not always measure semantic similarity, and gives analytical examples where it produces arbitrary or non-unique values.

419Safety
MM1: Multimodal LLM Pre-training

MM1: Multimodal LLM Pre-training

Apple's MM1 paper runs extensive ablations on multimodal LLM pretraining choices and releases a family of models up to 30B parameters that set competitive MLLM pretraining benchmarks.

420Training
Claude 3

Claude 3

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.

421Evaluation
Robust Evaluation of Reasoning

Robust Evaluation of Reasoning

The paper introduces functional benchmarks that parameterize reasoning problems so the same structural question can be re-instantiated with fresh surface forms, then uses them to expose a large "reasoning gap" in frontier LLMs.

422Reasoning
GaLore

GaLore

GaLore (Gradient Low-Rank Projection) reduces optimizer-state memory during LLM training while still permitting full-parameter updates, unlike LoRA-style adapters that restrict learning to a low-rank subspace.

423Memory
Can LLMs Reason and Plan?

Can LLMs Reason and Plan?

Kambhampati's position paper argues that what looks like reasoning and planning in LLMs is better understood as "universal approximate retrieval" powered by web-scale training.

424Reasoning
RAG for AI-Generated Content

RAG for AI-Generated Content

A survey that extends RAG beyond text, showing how retrieval augmentation is being applied across code, image, audio, video, and 3D generation.

425Retrieval
KnowAgent

KnowAgent

KnowAgent improves LLM-based planning agents by explicitly injecting action knowledge - what the actions are and how they relate - rather than letting the LLM invent its own action space at runtime.

426Agents
Sora Overview

Sora Overview

A comprehensive academic review of OpenAI's Sora, tracing the technical ingredients behind the text-to-video "world simulator" and the opportunities/limitations for the next wave of large vision models.

427Safety
SaulLM-7B: LLM for Law

SaulLM-7B: LLM for Law

SaulLM-7B is an open legal-domain LLM built on Mistral 7B and continually pretrained on 30B+ tokens of English legal text, with a companion instruction-tuning recipe.

428Training
Design2Code

Design2Code

Design2Code tackles the front-end engineering problem of turning a visual design into working HTML/CSS and gives the community both a benchmark and strong MLLM baselines.

429Evaluation
TripoSR

TripoSR

TripoSR is a transformer-based single-image 3D reconstruction model that returns a textured mesh in under 0.5 seconds, building on the LRM architecture with a stronger data and training pipeline.

430Training
Genie

Genie

DeepMind's Genie is an 11B-parameter foundation world model trained unsupervised on internet gameplay videos that generates action-controllable 2D worlds from a single image prompt.

431Agents
Mistral Large

Mistral Large

Mistral AI releases Mistral Large, its flagship closed-weight LLM positioned as the second-ranked API-accessible model behind GPT-4 at launch.

432Agents
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026