🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Training

Visual Instruction Tuning (LLaVA)

First page
Visual Instruction Tuning (LLaVA)
Paper summary

Uses language-only GPT-4 to generate multimodal instruction-following data.

Ask this paper

Key points
01

GPT-4-generated multimodal data: Bootstraps multimodal instruction data using only language-only GPT-4 given captions and bounding boxes - no direct visual access needed.

02

End-to-end training: Introduces LLaVA, an end-to-end trained large multimodal model combining CLIP vision encoder and Vicuna LLM.

03

Lightweight architecture: Simple projection layer between vision encoder and LLM - cheap and effective.

04

Open VLM revolution: LLaVA became the most influential open-source VLM architecture, spawning LLaVA-1.5, LLaVA-NeXT, and countless derivatives through 2024.

Every Monday
Get next week’s papers.
Subscribe on Substack