🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
618 papers · TrainingClear filters →
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

529Safety
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

530Safety
Radiology-Llama 2

Radiology-Llama 2

A Llama 2-based LLM specialized for radiology report generation.

531Training
Transformers as Support Vector Machines

Transformers as Support Vector Machines

A theoretical paper establishing a formal connection between self-attention optimization and hard-margin SVM problems.

532Training
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

533Reinforcement Learning
GPT Solves Math Problems Without a Calculator

GPT Solves Math Problems Without a Calculator

Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

534Reasoning
Explaining Grokking

Explaining Grokking

DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

535Safety
FLM-101B

FLM-101B

A 101B parameter open LLM trainable on a $100K budget through a growth-based training strategy.

536Training
LLaSM (Large Language and Speech Model)

LLaSM (Large Language and Speech Model)

A combined language-and-speech model trained with cross-modal conversational abilities.

537Multimodal
SAM-Med2D

SAM-Med2D

Adapts the Segment Anything Model (SAM) to 2D medical imaging through large-scale medical fine-tuning.

538Multimodal
Graph of Thoughts (GoT)

Graph of Thoughts (GoT)

Generalizes Chain-of-Thought and Tree-of-Thought by modeling LLM reasoning as an arbitrary graph.

539Reasoning
MVDream

MVDream

ByteDance's MVDream is a multi-view diffusion model that generates geometrically consistent images from multiple viewpoints given a text prompt.

540Multimodal
FaceChain

FaceChain

Alibaba's FaceChain is a personalized portrait generation framework that produces identity-preserving portraits from just a handful of input photos.

541Multimodal
Survey on Instruction Tuning for LLMs

Survey on Instruction Tuning for LLMs

A comprehensive survey of instruction tuning covering methodology, dataset construction, and applications.

542Training
Giraffe

Giraffe

A family of context-extended Llama and Llama 2 models, along with an empirical study of context-extension techniques.

543Memory
Prompt2Model

Prompt2Model

CMU's Prompt2Model automates the path from a natural-language task description to a deployable small special-purpose model.

544Agents
Humpback (Self-Alignment with Instruction Backtranslation)

Humpback (Self-Alignment with Instruction Backtranslation)

Meta's Humpback automatically generates instruction-tuning data by back-translating web text into plausible instructions.

545Safety
Platypus

Platypus

Platypus is a family of fine-tuned and merged LLMs that topped the Open LLM Leaderboard in August 2023.

546Training
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

547Efficiency
Shepherd

Shepherd

Meta's Shepherd is a 7B language model specifically tuned to critique model outputs and suggest refinements.

548Safety
Teach LLMs to Personalize

Teach LLMs to Personalize

A multitask-learning approach for personalized text generation without relying on predefined user attributes.

549Training
Political Biases in NLP Models

Political Biases in NLP Models

Develops methods to measure political and media biases in LLMs and their downstream effects.

550Safety
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

551Agents
SynJax

SynJax

DeepMind's SynJax is a JAX-based library for efficient vectorized inference in structured distributions.

552Training
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026