🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
301 papers · MemoryClear filters →
Effective Long-Context Scaling (Meta)

Effective Long-Context Scaling (Meta)

Meta proposes a 70B long-context LLM that surpasses GPT-3.5-turbo-16k on long-context benchmarks.

289Memory
LongLoRA

LongLoRA

An efficient LoRA-based fine-tuning recipe for extending LLM context windows without expensive full fine-tuning.

290Training
Giraffe

Giraffe

A family of context-extended Llama and Llama 2 models, along with an empirical study of context-extension techniques.

291Memory
L-Eval

L-Eval

A standardized evaluation suite for long-context language models.

292Evaluation
FlashAttention-2

FlashAttention-2

Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

293Efficiency
Retentive Network (RetNet)

Retentive Network (RetNet)

Microsoft's proposed foundation architecture aiming to replace Transformer attention for LLMs.

294Architecture
LongLLaMA

LongLLaMA

Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

295Memory
How Language Models Use Long Contexts (Lost-in-the-Middle)

How Language Models Use Long Contexts (Lost-in-the-Middle)

Shows LLM performance drops when relevant information is in the middle of a long context.

296Memory
Extending Context Window of LLMs (PI)

Extending Context Window of LLMs (PI)

Position Interpolation extends LLaMA's context to 32K with minimal fine-tuning (within 1000 steps).

297Memory
Augmenting LLMs with Long-term Memory (LongMem)

Augmenting LLMs with Long-term Memory (LongMem)

Enables LLMs to memorize long history via memory-augmented adaptation.

298Memory
Augmenting LLMs with Databases (ChatDB)

Augmenting LLMs with Databases (ChatDB)

Combines an LLM with SQL databases as a symbolic memory framework.

299Memory
Unlimiformer

Unlimiformer

Long-range Transformers with unlimited length input via external datastores.

300Memory
Scaling Transformer to 1M tokens with RMT

Scaling Transformer to 1M tokens with RMT

Recurrent Memory Transformer extends BERT's effective context to 2M tokens.

301Memory
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026