
CM3Leon
Meta's retrieval-augmented multi-modal language model that generates both text and images.

Claude 2
Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

Secrets of RLHF in LLMs
A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.

LongLLaMA
Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

Patch n' Pack: NaViT
A vision transformer handling any aspect ratio and resolution through sequence packing.

LLMs as General Pattern Machines
Demonstrates LLMs serve as general sequence modelers without additional training.

HyperDreamBooth
A smaller, faster, and more efficient version of DreamBooth for personalizing text-to-image models.

Teaching Arithmetic to Small Transformers
Trains small transformers on chain-of-thought style data for arithmetic with large gains.

AnimateDiff
Animates frozen text-to-image diffusion models via a plug-in motion modeling module.

Generative Pretraining in Multimodality (Emu)
A transformer-based multimodal foundation model for generating images and text.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack