🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
LLMs Represent Space and Time

LLMs Represent Space and Time

MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

121Reasoning
Retrieval Meets Long-Context LLMs

Retrieval Meets Long-Context LLMs

NVIDIA's study comparing RAG and long-context LLMs, with the punchline that the two are complementary rather than substitutes.

122Retrieval
StreamingLLM

StreamingLLM

MIT's StreamingLLM enables efficient streaming inference by preserving "attention sinks" - early-sequence tokens that most attention mass flows to.

123Efficiency
Neural Developmental Programs (NDPs)

Neural Developmental Programs (NDPs)

Proposes neural networks that self-assemble through a developmental process inspired by biological embryonic development.

124Architecture
The Dawn of LMMs (GPT-4V Deep Dive)

The Dawn of LMMs (GPT-4V Deep Dive)

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.

125Multimodal
Training LLMs with Pause Tokens

Training LLMs with Pause Tokens

CMU shows that adding a learnable `<pause>` token during both pretraining and fine-tuning gives the model extra "thinking time" and improves reasoning.

126Training
Self-Taught Optimizer (STOP)

Self-Taught Optimizer (STOP)

Proposes recursively self-improving code generation where an LLM-scaffolded program improves itself.

127Code
RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

Meta's RA-DIT is a lightweight recipe that retrofits LLMs with retrieval capabilities through dual fine-tuning.

128Retrieval
KOSMOS-G

KOSMOS-G

Microsoft's KOSMOS-G extends zero-shot image generation to multi-image vision-language input.

129Multimodal
Analogical Prompting

Analogical Prompting

Google's Analogical Prompting guides LLM reasoning by having the model self-generate relevant exemplars on the fly.

130Reasoning
The Reversal Curse

The Reversal Curse

Finds that LLMs trained on "A is B" fail to generalize to "B is A" - a surprisingly deep failure of learning.

131Training
Effective Long-Context Scaling (Meta)

Effective Long-Context Scaling (Meta)

Meta proposes a 70B long-context LLM that surpasses GPT-3.5-turbo-16k on long-context benchmarks.

132Memory
Graph Neural Prompting (GNP)

Graph Neural Prompting (GNP)

A plug-and-play method that injects knowledge-graph information into frozen pretrained LLMs.

133Training
Vision Transformers Need Registers

Vision Transformers Need Registers

Meta researchers identify artifact tokens in ViT feature maps and propose a trivial fix: add dedicated register tokens.

134Training
Boolformer

Boolformer

The first Transformer trained to perform end-to-end symbolic regression of Boolean functions.

135Reasoning
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

136Reinforcement Learning
LLM Alignment Survey

LLM Alignment Survey

A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

137Safety
Qwen

Qwen

Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

138Training
MentaLLaMA

MentaLLaMA

An open-source LLM family specialized for interpretable mental-health analysis on social media.

139Safety
Logical Chain-of-Thought (LogiCoT)

Logical Chain-of-Thought (LogiCoT)

A neurosymbolic framework that verifies and revises zero-shot CoT reasoning using symbolic-logic principles.

140Reasoning
AlphaMissense

AlphaMissense

DeepMind's AlphaMissense is an AI model that classifies missense genetic variants as pathogenic or benign at genome scale.

141Architecture
Chain-of-Verification (CoVe)

Chain-of-Verification (CoVe)

Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

142Safety
Contrastive Decoding for Reasoning

Contrastive Decoding for Reasoning

Shows that contrastive decoding, a simple inference-time technique, substantially improves reasoning in large LLMs.

143Reasoning
LongLoRA

LongLoRA

An efficient LoRA-based fine-tuning recipe for extending LLM context windows without expensive full fine-tuning.

144Training
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026