🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
CogAgent

CogAgent

Tsinghua's CogAgent is an 18B-parameter visual-language model purpose-built for GUI understanding and navigation, with unusually high input resolution.

01Evaluation
From Gemini to Q-Star

From Gemini to Q-Star

A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

02Multimodal
PromptBench

PromptBench

A unified library for comprehensive evaluation and analysis of LLMs that consolidates multiple evaluation concerns under one roof.

03Evaluation
Exploiting Novel GPT-4 APIs

Exploiting Novel GPT-4 APIs

A red-team study of three newer GPT-4 API surfaces - fine-tuning, function calling, and knowledge retrieval - that reveals each introduces new attack vectors.

04Training
Fact Recalling in LLMs

Fact Recalling in LLMs

A mechanistic-interpretability study showing that early MLP layers function as a lookup table for factual recall.

05Safety
Generative AI for Math (OpenWebMath / MathPile)

Generative AI for Math (OpenWebMath / MathPile)

Releases a diverse, high-quality math-centric corpus of ~9.5B tokens designed for training math-capable foundation models.

06Reasoning
Principled Instructions Are All You Need

Principled Instructions Are All You Need

Distills effective LLM prompting into 26 guiding principles and validates them across multiple model families.

07Training
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

08Reasoning
LLaRA

LLaRA

LLaRA adapts a decoder-only LLM for dense retrieval via two tailored pretext tasks that leverage text embeddings from the LLM itself.

09Retrieval
Gemini vs GPT-4V

Gemini vs GPT-4V

A qualitative side-by-side comparison of Gemini and GPT-4V across vision-language tasks, documenting systematic behavioral differences.

10Multimodal
Gemini's Language Abilities

Gemini's Language Abilities

CMU's impartial, reproducible evaluation of Gemini Pro against GPT and Mixtral across standard LLM benchmarks.

11Evaluation
PowerInfer

PowerInfer

A high-speed LLM inference engine for consumer GPUs that exploits sparse neuron activation patterns to run large models on commodity hardware.

12Efficiency
Antibiotic Discovery with Graph Deep Learning (Nature)

Antibiotic Discovery with Graph Deep Learning (Nature)

MIT researchers use explainable graph neural networks to discover a new structural class of antibiotics.

13Training
VideoPoet

VideoPoet

Google Research's VideoPoet is a large language model for zero-shot video generation that treats video as just another token stream.

14Multimodal
AppAgent

AppAgent

Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

15Multimodal
LLM in a Flash

LLM in a Flash

Apple researchers show how to run LLMs larger than available DRAM by streaming weights from flash storage on demand.

16Memory
ReST Meets ReAct

ReST Meets ReAct

Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

17Agents
Adversarial Attacks on GPT-4

Adversarial Attacks on GPT-4

Demonstrates that a trivially simple random-search procedure can jailbreak GPT-4 with high reliability.

18Safety
RAG for LLMs

RAG for LLMs

A broad survey of Retrieval-Augmented Generation research, organizing the rapidly growing literature into a coherent map.

19Retrieval
BabyLLM Challenge Findings

BabyLLM Challenge Findings

Reports results from a challenge on sample-efficient pretraining using a developmentally plausible corpus.

20Training
FunSearch

FunSearch

DeepMind's FunSearch uses LLMs as a mutation operator in an evolutionary loop to discover genuinely new mathematical knowledge.

21Safety
Weak-to-Strong Generalization

Weak-to-Strong Generalization

OpenAI's superalignment team shows that weak supervisors can still elicit capabilities from much stronger models - a first empirical signal for scalable oversight.

22Training
Audiobox

Audiobox

Meta's Audiobox is a unified flow-matching audio model that generates speech, sound effects, and music from natural-language and example prompts.

23Multimodal
Mathematical LLMs Survey

Mathematical LLMs Survey

A survey on the progress of LLMs on mathematical reasoning tasks, covering methods, benchmarks, and open problems.

24Reasoning
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026