🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Multimodal

MoE-LLaVA

First page
MoE-LLaVA
Paper summary

MoE-LLaVA applies Mixture-of-Experts tuning to the LLaVA vision-language architecture, getting a sparse model with dramatically fewer active parameters at the same compute cost.

Ask this paper

Key points
01

MoE for VLMs: Replaces the dense feed-forward layers in LLaVA's decoder with sparse MoE layers and introduces a three-stage training recipe that stabilizes the typically fragile MoE + multimodal combo.

02

Efficient inference: Only a fraction of experts are active per token, giving the model effective sparsity while keeping FLOPs per forward pass similar to a smaller dense baseline.

03

Matches larger dense VLMs: A MoE-LLaVA with 3B activated parameters matches or beats 7B dense VLMs on standard multimodal benchmarks, closing the parameter/performance gap.

04

Mitigates sparsity degradation: The auxiliary balancing losses and staged training schedule address the usual MoE pitfalls of routing collapse and modality interference.

Every Monday
Get next week’s papers.
Subscribe on Substack