🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
This week · 10 papersView the full issue →
ImageBind-LLM

ImageBind-LLM

Shanghai AI Lab's ImageBind-LLM brings six-modality understanding to LLMs via the ImageBind joint embedding space.

02Multimodal
Explaining Grokking

Explaining Grokking

DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

03Safety
Overview of AI Deception

Overview of AI Deception

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

04Safety
FLM-101B

FLM-101B

A 101B parameter open LLM trainable on a $100K budget through a growth-based training strategy.

05Training
Cognitive Architectures for Language Agents (CoALA)

Cognitive Architectures for Language Agents (CoALA)

Princeton proposes CoALA, a systematic framework for understanding and building language agents.

06Agents
Q-Transformer

Q-Transformer

Google's Q-Transformer is a scalable RL method for training multi-task robotic policies from large offline datasets.

07Robotics
LLaSM (Large Language and Speech Model)

LLaSM (Large Language and Speech Model)

A combined language-and-speech model trained with cross-modal conversational abilities.

08Multimodal
SAM-Med2D

SAM-Med2D

Adapts the Segment Anything Model (SAM) to 2D medical imaging through large-scale medical fine-tuning.

09Training
Vector Search with OpenAI Embeddings

Vector Search with OpenAI Embeddings

Argues, via empirical analysis, that dedicated vector databases aren't necessarily required for modern AI-stack search applications.

10Retrieval
Graph of Thoughts (GoT)

Graph of Thoughts (GoT)

Generalizes Chain-of-Thought and Tree-of-Thought by modeling LLM reasoning as an arbitrary graph.

11Reasoning
MVDream

MVDream

ByteDance's MVDream is a multi-view diffusion model that generates geometrically consistent images from multiple viewpoints given a text prompt.

12Multimodal
Nougat

Nougat

Meta's Nougat is a visual transformer for "Neural Optical Understanding for Academic documents" that converts PDFs to LaTeX/Markdown.

13Multimodal
FacTool

FacTool

A tool-augmented framework for detecting factual errors in LLM-generated text.

14Retrieval
AnomalyGPT

AnomalyGPT

Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

15Data
FaceChain

FaceChain

Alibaba's FaceChain is a personalized portrait generation framework that produces identity-preserving portraits from just a handful of input photos.

16Multimodal
Qwen-VL

Qwen-VL

Alibaba's Qwen-VL is a large-scale vision-language model family with strong performance across captioning, VQA, and visual localization.

17Multimodal
Code Llama

Code Llama

Meta releases Code Llama, a family of code-specialized LLMs built on top of Llama 2.

18Code
Survey on Instruction Tuning for LLMs

Survey on Instruction Tuning for LLMs

A comprehensive survey of instruction tuning covering methodology, dataset construction, and applications.

19Training
SeamlessM4T

SeamlessM4T

Meta's SeamlessM4T is a unified multilingual and multimodal machine-translation system that handles five translation tasks in one model.

20Multimodal
LLMs for Illicit Purposes

LLMs for Illicit Purposes

A survey cataloguing threats and vulnerabilities arising from LLM deployment.

21Safety
Giraffe

Giraffe

A family of context-extended Llama and Llama 2 models, along with an empirical study of context-extension techniques.

22Memory
IT3D

IT3D

Improves Text-to-3D generation by leveraging explicitly synthesized multi-view images in the training loop.

23Multimodal
LLM-Based Autonomous Agents Survey

LLM-Based Autonomous Agents Survey

A comprehensive survey of LLM-based autonomous agents covering construction and applications.

24Agents
Prompt2Model

Prompt2Model

CMU's Prompt2Model automates the path from a natural-language task description to a deployable small special-purpose model.

25Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026