AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

DINOv2
Meta's self-supervised vision foundation model producing robust features without labels.

Learning to Compress Prompts with Gist Tokens
Trains LMs to compress prompts into reusable "gist" tokens.

Scaling Biomolecular Simulations with Equivariant Models
A framework for large-scale biomolecular simulation using equivariant deep learning.

Evaluating Verifiability in Generative Search Engines
Audits popular generative search engines for citation accuracy.

Generative Disco: Text-to-Video Generation for Music Visualization
An LLM + T2I system for music visualization.

Architectures of Topological Deep Learning: A Survey on Topological Neural Networks
A comprehensive survey on topological neural networks.

Visual Instruction Tuning (LLaVA)
Uses language-only GPT-4 to generate multimodal instruction-following data.

ChatGPT: Applications, Opportunities, and Threats
A comprehensive overview of ChatGPT's applications and risks.

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
A framework inferring tool sequences for compositional reasoning.

Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
High-resolution video synthesis with latent diffusion.

Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields
Combines mip-NeRF 360 with grid-based models for 22x faster training.

Generative Agents: Interactive Simulacra of Human Behavior
Stanford/Google's landmark paper on LLM-powered social simulations.

Emergent Autonomous Scientific Research Capabilities of LLMs
An agent combining LLMs for autonomous scientific experiments.

Automatic Gradient Descent: Deep Learning without Hyperparameters
A hyperparameter-free first-order optimizer that leverages architecture.

ChemCrow: Augmenting LLMs with Chemistry Tools
An LLM chemistry agent with 13 expert-designed tools.

One Small Step for Generative AI, One Giant Leap for AGI
A complete survey on ChatGPT and GPT-4.

OpenAGI: When LLM Meets Domain Experts
An open-source research platform for LLM agents manipulating domain expert models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
A benchmark using real human standardized exams.

Teaching Large Language Models to Self-Debug
Teaches LLMs to debug their own code via few-shot demonstrations.

Segment Everything Everywhere All at Once (SEEM)
A promptable, interactive segmentation model.

Segment Anything (SAM)
Meta's foundational model for image segmentation with massive training data release.

Instruction Tuning with GPT-4
Uses GPT-4 to generate instruction-following data for LLM fine-tuning.

Eight Things to Know about Large Language Models
Sam Bowman's influential primer on key LLM considerations.

A Survey of Large Language Models
A 50-page comprehensive survey on LLMs.