🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
860 papers · EvaluationClear filters →
SAM-Med2D

SAM-Med2D

Adapts the Segment Anything Model (SAM) to 2D medical imaging through large-scale medical fine-tuning.

769Multimodal
Vector Search with OpenAI Embeddings

Vector Search with OpenAI Embeddings

Argues, via empirical analysis, that dedicated vector databases aren't necessarily required for modern AI-stack search applications.

770Retrieval
AnomalyGPT

AnomalyGPT

Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

771Data
Code Llama

Code Llama

Meta releases Code Llama, a family of code-specialized LLMs built on top of Llama 2.

772Code
Survey on Instruction Tuning for LLMs

Survey on Instruction Tuning for LLMs

A comprehensive survey of instruction tuning covering methodology, dataset construction, and applications.

773Training
SeamlessM4T

SeamlessM4T

Meta's SeamlessM4T is a unified multilingual and multimodal machine-translation system that handles five translation tasks in one model.

774Multimodal
LLMs for Illicit Purposes

LLMs for Illicit Purposes

A survey cataloguing threats and vulnerabilities arising from LLM deployment.

775Safety
LLM-Based Autonomous Agents Survey

LLM-Based Autonomous Agents Survey

A comprehensive survey of LLM-based autonomous agents covering construction and applications.

776Agents
Prompt2Model

Prompt2Model

CMU's Prompt2Model automates the path from a natural-language task description to a deployable small special-purpose model.

777Agents
LegalBench

LegalBench

A collaboratively constructed benchmark for measuring legal reasoning in LLMs.

778Evaluation
Language to Rewards for Robotic Skill Synthesis

Language to Rewards for Robotic Skill Synthesis

Google's Language-to-Rewards uses LLMs to define reward parameters for robotic RL.

779Robotics
Humpback (Self-Alignment with Instruction Backtranslation)

Humpback (Self-Alignment with Instruction Backtranslation)

Meta's Humpback automatically generates instruction-tuning data by back-translating web text into plausible instructions.

780Safety
Platypus

Platypus

Platypus is a family of fine-tuned and merged LLMs that topped the Open LLM Leaderboard in August 2023.

781Training
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

782Efficiency
GEARS

GEARS

Stanford's GEARS predicts cellular responses to genetic perturbation using deep learning + a gene-relationship knowledge graph.

783Evaluation
OctoPack

OctoPack

Hugging Face releases OctoPack, a 4TB dataset of Git commits across 350 programming languages for instruction-tuning code LLMs.

784Data
Bayesian Flow Networks (BFN)

Bayesian Flow Networks (BFN)

Introduces a new class of generative models that combine Bayesian inference with deep learning.

785Architecture
D-Bot (LLMs as Database Administrators)

D-Bot (LLMs as Database Administrators)

Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

786Agents
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

787Agents
PUG (Photorealistic Unreal Graphics)

PUG (Photorealistic Unreal Graphics)

Meta's PUG uses Unreal Engine to generate photorealistic, semantically controllable synthetic datasets for vision research.

788Data
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

789Safety
Open Problems and Limitations of RLHF

Open Problems and Limitations of RLHF

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

790Reinforcement Learning
Med-Flamingo

Med-Flamingo

Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

791Multimodal
ToolLLM

ToolLLM

Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

792Agents
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026