Visual Instruction Tuning (LLaVA)
First page

Paper summary
Uses language-only GPT-4 to generate multimodal instruction-following data.
Ask this paper
01
GPT-4-generated multimodal data: Bootstraps multimodal instruction data using only language-only GPT-4 given captions and bounding boxes - no direct visual access needed.
02
End-to-end training: Introduces LLaVA, an end-to-end trained large multimodal model combining CLIP vision encoder and Vicuna LLM.
03
Lightweight architecture: Simple projection layer between vision encoder and LLM - cheap and effective.
04
Open VLM revolution: LLaVA became the most influential open-source VLM architecture, spawning LLaVA-1.5, LLaVA-NeXT, and countless derivatives through 2024.