🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
860 papers · EvaluationClear filters →
OpenFlamingo

OpenFlamingo

An open-source family of autoregressive vision-language models spanning 3B to 9B parameters.

793Multimodal
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

794Safety
Med-PaLM Multimodal

Med-PaLM Multimodal

Introduces a generalist biomedical AI system and a new multimodal biomedical benchmark with 14 tasks.

795Multimodal
Foundation Models in Vision

Foundation Models in Vision

A comprehensive survey on foundational models for computer vision and their open research directions.

796Multimodal
L-Eval

L-Eval

A standardized evaluation suite for long-context language models.

797Evaluation
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

798Safety
FacTool

FacTool

A task- and domain-agnostic framework for factuality detection of LLM-generated text.

799Evaluation
How is ChatGPT's Behavior Changing Over Time?

How is ChatGPT's Behavior Changing Over Time?

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

800Evaluation
Challenges & Application of LLMs

Challenges & Application of LLMs

A comprehensive enumeration of open challenges and application domains for LLMs.

801Safety
Retrieve In-Context Examples for LLMs

Retrieve In-Context Examples for LLMs

A framework to iteratively train dense retrievers that identify high-quality in-context examples.

802Retrieval
FLASK

FLASK

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.

803Evaluation
CM3Leon

CM3Leon

Meta's retrieval-augmented multi-modal language model that generates both text and images.

804Multimodal
Claude 2

Claude 2

Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

805Safety
LongLLaMA

LongLLaMA

Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

806Memory
LLMs as General Pattern Machines

LLMs as General Pattern Machines

Demonstrates LLMs serve as general sequence modelers without additional training.

807Reasoning
A Survey on Evaluation of LLMs

A Survey on Evaluation of LLMs

A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.

808Evaluation
How Language Models Use Long Contexts (Lost-in-the-Middle)

How Language Models Use Long Contexts (Lost-in-the-Middle)

Shows LLM performance drops when relevant information is in the middle of a long context.

809Memory
LLMs as Effective Text Rankers

LLMs as Effective Text Rankers

A prompting technique that enables open-source LLMs to perform SOTA text ranking.

810Retrieval
Elastic Decision Transformer

Elastic Decision Transformer

An advance over Decision Transformers that enables trajectory stitching at inference time.

811Reinforcement Learning
InterCode

InterCode

A framework treating interactive coding as a reinforcement learning environment.

812Reinforcement Learning
LeanDojo

LeanDojo

An open-source Lean playground consisting of toolkits, data, models, and benchmarks for theorem proving.

813Reasoning
Extending Context Window of LLMs (PI)

Extending Context Window of LLMs (PI)

Position Interpolation extends LLaMA's context to 32K with minimal fine-tuning (within 1000 steps).

814Memory
Understanding Theory-of-Mind in LLMs with LLMs

Understanding Theory-of-Mind in LLMs with LLMs

A framework for procedurally generating ToM evaluations using LLMs themselves.

815Evaluation
Evaluations with No Labels

Evaluations with No Labels

Self-supervised evaluation of LLMs via sensitivity/invariance to input transformations.

816Evaluation
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026