MM1: Multimodal LLM Pre-training
Free while signed in. Answers cite the passages they came from.

Apple's MM1 paper runs extensive ablations on multimodal LLM pretraining choices and releases a family of models up to 30B parameters that set competitive MLLM pretraining benchmarks.
Data mixture is key: Image-caption, interleaved image-text, and text-only data each play distinct roles; the right mix is essential for few-shot and zero-shot performance.
Image encoder dominates: Encoder choice and input resolution matter a lot, while the vision-language connector design turns out to be comparatively unimportant.
30B dense + MoE: The MM1 family includes both dense and mixture-of-experts variants up to 30B parameters, with strong SoTA pretraining metrics.
Emergent capabilities: After pretraining, MM1 exhibits multi-image reasoning and supports few-shot chain-of-thought prompting - capabilities that earlier smaller MLLMs typically lacked.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack