Generative Pretraining in Multimodality (Emu)
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsA transformer-based multimodal foundation model for generating images and text.
01
Unified pretraining: Pretrains on mixed image-text sequences to generate either modality in multimodal context.
02
Instruction tuning for assistants: Combines generative pretraining with instruction tuning to produce performant multimodal assistants.
03
In-context multimodal: Supports in-context learning across images and text, enabling few-shot multimodal tasks.
04
Multi-modal assistants: Part of the 2023 push (alongside LLaVA, MiniGPT-4) that established the pattern of visual-instruction-tuned assistants.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack