
CM3Leon
Meta's retrieval-augmented multi-modal language model that generates both text and images.

Claude 2
Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

Secrets of RLHF in LLMs
A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.

LongLLaMA
Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

Patch n' Pack: NaViT
A vision transformer handling any aspect ratio and resolution through sequence packing.

LLMs as General Pattern Machines
Demonstrates LLMs serve as general sequence modelers without additional training.

HyperDreamBooth
A smaller, faster, and more efficient version of DreamBooth for personalizing text-to-image models.

Teaching Arithmetic to Small Transformers
Trains small transformers on chain-of-thought style data for arithmetic with large gains.

AnimateDiff
Animates frozen text-to-image diffusion models via a plug-in motion modeling module.

Generative Pretraining in Multimodality (Emu)
A transformer-based multimodal foundation model for generating images and text.