🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Training

Advances in Multimodal LLMs

First page
Advances in Multimodal LLMs
Paper summary

A comprehensive survey mapping design choices for architecture and training pipeline around multimodal large language models (MLLMs).

Ask this paper

Key points
01

Architecture taxonomy: Organizes MLLMs by modality encoders, LLM backbones, and cross-modal connectors (Q-Former, linear, cross-attention, etc.), clarifying which design axes matter most.

02

Training pipeline: Walks through modality alignment pretraining, visual instruction tuning, and RLHF-style preference optimization as the dominant recipe across recent MLLMs.

03

Applications and evaluation: Catalogs tasks (VQA, captioning, grounding, video, document understanding) alongside the benchmarks commonly used to evaluate them.

04

Open challenges: Identifies hallucination, long-video understanding, efficient training, and tool-augmented multimodality as the most pressing research directions.

Every Monday
Get next week’s papers.
Subscribe on Substack