
DINOv2
Meta's self-supervised vision foundation model producing robust features without labels.

Learning to Compress Prompts with Gist Tokens
Trains LMs to compress prompts into reusable "gist" tokens.

Scaling Biomolecular Simulations with Equivariant Models
A framework for large-scale biomolecular simulation using equivariant deep learning.

Evaluating Verifiability in Generative Search Engines
Audits popular generative search engines for citation accuracy.

Generative Disco: Text-to-Video Generation for Music Visualization
An LLM + T2I system for music visualization.

Architectures of Topological Deep Learning: A Survey on Topological Neural Networks
A comprehensive survey on topological neural networks.

Visual Instruction Tuning (LLaVA)
Uses language-only GPT-4 to generate multimodal instruction-following data.

ChatGPT: Applications, Opportunities, and Threats
A comprehensive overview of ChatGPT's applications and risks.

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
A framework inferring tool sequences for compositional reasoning.

Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
High-resolution video synthesis with latent diffusion.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack