🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training · Multimodal

MM1: Multimodal LLM Pre-training

Free while signed in. Answers cite the passages they came from.

First page
MM1: Multimodal LLM Pre-training
The curator’s take

Apple's MM1 paper runs extensive ablations on multimodal LLM pretraining choices and releases a family of models up to 30B parameters that set competitive MLLM pretraining benchmarks.

Key points
01

Data mixture is key: Image-caption, interleaved image-text, and text-only data each play distinct roles; the right mix is essential for few-shot and zero-shot performance.

02

Image encoder dominates: Encoder choice and input resolution matter a lot, while the vision-language connector design turns out to be comparatively unimportant.

03

30B dense + MoE: The MM1 family includes both dense and mixture-of-experts variants up to 30B parameters, with strong SoTA pretraining metrics.

04

Emergent capabilities: After pretraining, MM1 exhibits multi-image reasoning and supports few-shot chain-of-thought prompting - capabilities that earlier smaller MLLMs typically lacked.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack