🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Transfusion

First page
Transfusion
Paper summary

presents a training recipe to train multi-modal models over discrete and continuous data; combines next token prediction with diffusion to train transformer models over mixed-modality sequences; shows that it’s possible to scale from 7B parameter models to 2T multi-modal tokens that can compete in performance with similar scale diffusion and language models.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack