Generative Pretraining in Multimodality (Emu)
First page

Paper summary
A transformer-based multimodal foundation model for generating images and text.
Ask this paper
01
Unified pretraining: Pretrains on mixed image-text sequences to generate either modality in multimodal context.
02
Instruction tuning for assistants: Combines generative pretraining with instruction tuning to produce performant multimodal assistants.
03
In-context multimodal: Supports in-context learning across images and text, enabling few-shot multimodal tasks.
04
Multi-modal assistants: Part of the 2023 push (alongside LLaVA, MiniGPT-4) that established the pattern of visual-instruction-tuned assistants.