AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

CogAgent
Tsinghua's CogAgent is an 18B-parameter visual-language model purpose-built for GUI understanding and navigation, with unusually high input resolution.

From Gemini to Q-Star
A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

Survey of Reasoning with Foundation Models
A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

Gemini vs GPT-4V
A qualitative side-by-side comparison of Gemini and GPT-4V across vision-language tasks, documenting systematic behavioral differences.

VideoPoet
Google Research's VideoPoet is a large language model for zero-shot video generation that treats video as just another token stream.

AppAgent
Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

RAG for LLMs
A broad survey of Retrieval-Augmented Generation research, organizing the rapidly growing literature into a coherent map.

Audiobox
Meta's Audiobox is a unified flow-matching audio model that generates speech, sound effects, and music from natural-language and example prompts.

Gemini 1.0
Google launches Gemini 1.0, a multimodal family natively designed to reason across text, images, video, audio, and code from the ground up.

Adversarial Diffusion Distillation (SDXL Turbo)
Stability AI's ADD trains a student diffusion model that produces high-quality images in just 1-4 sampling steps.

Seamless
Meta's Seamless is a family of models for end-to-end expressive, streaming cross-lingual speech communication.

UniIR
UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

Mirasol3B
Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

GAIA
Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

Emu Video and Emu Edit
Meta releases Emu Video and Emu Edit, a pair of diffusion models targeting controlled text-to-video generation and instruction-based image editing.

JARVIS-1
An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

MusicGen
Meta's MusicGen is a single-stage transformer LLM for music generation that operates over compressed discrete audio tokens.

MetNet-3
Google's MetNet-3 is a state-of-the-art neural weather model extending lead time and variable coverage well beyond prior observation-based models.

Matryoshka Diffusion Models
Apple introduces an end-to-end framework for high-resolution image and video synthesis that denoises across multiple resolutions jointly.

Spectron
Google's Spectron is a spoken-language model trained end-to-end on raw spectrograms rather than text or discrete audio tokens.

CommonCanvas
Releases CommonCanvas, a text-to-image dataset composed entirely of Creative-Commons-licensed images.

Video Language Planning
Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

UniSim (Universal Simulator)
Google's UniSim learns a universal generative simulator of real-world interactions from diverse video + action data.

The Dawn of LMMs (GPT-4V Deep Dive)
Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.