🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

193Efficiency
GEARS

GEARS

Stanford's GEARS predicts cellular responses to genetic perturbation using deep learning + a gene-relationship knowledge graph.

194Evaluation
Shepherd

Shepherd

Meta's Shepherd is a 7B language model specifically tuned to critique model outputs and suggest refinements.

195Safety
GPT-4 Code Interpreter for Math

GPT-4 Code Interpreter for Math

A zero-shot prompting technique for GPT-4 Code Interpreter that dramatically boosts math-reasoning accuracy via code self-verification.

196Reasoning
Teach LLMs to Personalize

Teach LLMs to Personalize

A multitask-learning approach for personalized text generation without relying on predefined user attributes.

197Training
OctoPack

OctoPack

Hugging Face releases OctoPack, a 4TB dataset of Git commits across 350 programming languages for instruction-tuning code LLMs.

198Data
Outlines (Efficient Guided Generation)

Outlines (Efficient Guided Generation)

A library for guided LLM text generation that enforces structural constraints with minimal overhead.

199Efficiency
Bayesian Flow Networks (BFN)

Bayesian Flow Networks (BFN)

Introduces a new class of generative models that combine Bayesian inference with deep learning.

200Architecture
D-Bot (LLMs as Database Administrators)

D-Bot (LLMs as Database Administrators)

Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

201Agents
Political Biases in NLP Models

Political Biases in NLP Models

Develops methods to measure political and media biases in LLMs and their downstream effects.

202Safety
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

203Agents
Studying LLM Generalization with Influence Functions

Studying LLM Generalization with Influence Functions

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

204Safety
NeuroImagen

NeuroImagen

Reconstructs visual stimuli images from EEG signals using latent diffusion, opening new windows into visually-evoked brain activity.

205Multimodal
SynJax

SynJax

DeepMind's SynJax is a JAX-based library for efficient vectorized inference in structured distributions.

206Training
Synthetic Data Reduces Sycophancy

Synthetic Data Reduces Sycophancy

Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

207Data
PUG (Photorealistic Unreal Graphics)

PUG (Photorealistic Unreal Graphics)

Meta's PUG uses Unreal Engine to generate photorealistic, semantically controllable synthetic datasets for vision research.

208Data
LLMs for HVAC Control

LLMs for HVAC Control

Microsoft applies LLMs to industrial control tasks (HVAC for buildings), comparing against RL baselines.

209Agents
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

210Safety
Open Problems and Limitations of RLHF

Open Problems and Limitations of RLHF

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

211Reinforcement Learning
Med-Flamingo

Med-Flamingo

Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

212Multimodal
ToolLLM

ToolLLM

Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

213Agents
Skeleton-of-Thought (SoT)

Skeleton-of-Thought (SoT)

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

214Reasoning
MetaGPT

MetaGPT

MetaGPT is a multi-agent framework that encodes standardized operating procedures (SOPs) for complex problem solving.

215Agents
OpenFlamingo

OpenFlamingo

An open-source family of autoregressive vision-language models spanning 3B to 9B parameters.

216Multimodal
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026