🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
355 papers · MemoryClear filters →
FlashAttention-2

FlashAttention-2

Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

337Efficiency
Retentive Network (RetNet)

Retentive Network (RetNet)

Microsoft's proposed foundation architecture aiming to replace Transformer attention for LLMs.

338Architecture
Claude 2

Claude 2

Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

339Safety
LongLLaMA

LongLLaMA

Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

340Memory
How Language Models Use Long Contexts (Lost-in-the-Middle)

How Language Models Use Long Contexts (Lost-in-the-Middle)

Shows LLM performance drops when relevant information is in the middle of a long context.

341Memory
Scaling Transformer to 1 Billion Tokens (LongNet)

Scaling Transformer to 1 Billion Tokens (LongNet)

Microsoft's Transformer variant scaling sequence length past 1B tokens.

342Memory
Extending Context Window of LLMs (PI)

Extending Context Window of LLMs (PI)

Position Interpolation extends LLaMA's context to 32K with minimal fine-tuning (within 1000 steps).

343Memory
Long-range Language Modeling with Self-Retrieval

Long-range Language Modeling with Self-Retrieval

Jointly trains a retrieval-augmented LM from scratch for long-range modeling.

344Retrieval
LOMO

LOMO

A memory-efficient optimizer that combines gradient computation and parameter update in one step.

345Efficiency
Augmenting LLMs with Long-term Memory (LongMem)

Augmenting LLMs with Long-term Memory (LongMem)

Enables LLMs to memorize long history via memory-augmented adaptation.

346Memory
Augmenting LLMs with Databases (ChatDB)

Augmenting LLMs with Databases (ChatDB)

Combines an LLM with SQL databases as a symbolic memory framework.

347Memory
No Positional Encodings (NoPE)

No Positional Encodings (NoPE)

Shows explicit position embeddings aren't essential for decoder-only Transformers.

348Architecture
Fine-Tuning Language Models with Just Forward Passes (MeZO)

Fine-Tuning Language Models with Just Forward Passes (MeZO)

A memory-efficient zeroth-order optimizer for LLM fine-tuning.

349Training
QLoRA

QLoRA

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.

350Training
Reinventing RNNs for the Transformer Era (RWKV)

Reinventing RNNs for the Transformer Era (RWKV)

Combines parallelizable training of Transformers with efficient RNN inference.

351Architecture
StarCoder

StarCoder

An open-access 15.5B code LLM with 8K context and 80+ programming languages.

352Code
Learning to Reason and Memorize with Self-Notes

Learning to Reason and Memorize with Self-Notes

LLMs that deviate from input to explicitly "think" and memorize.

353Reasoning
Scaling Transformer to 1M tokens with RMT

Scaling Transformer to 1M tokens with RMT

Recurrent Memory Transformer extends BERT's effective context to 2M tokens.

354Memory
Generative Agents: Interactive Simulacra of Human Behavior

Generative Agents: Interactive Simulacra of Human Behavior

Stanford/Google's landmark paper on LLM-powered social simulations.

355Agents
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026