🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Multimodal · Training

Visual Instruction Tuning (LLaVA)

Free while signed in. Answers cite the passages they came from.

First page
Visual Instruction Tuning (LLaVA)
The curator’s take

Uses language-only GPT-4 to generate multimodal instruction-following data.

Key points
01

GPT-4-generated multimodal data: Bootstraps multimodal instruction data using only language-only GPT-4 given captions and bounding boxes - no direct visual access needed.

02

End-to-end training: Introduces LLaVA, an end-to-end trained large multimodal model combining CLIP vision encoder and Vicuna LLM.

03

Lightweight architecture: Simple projection layer between vision encoder and LLM - cheap and effective.

04

Open VLM revolution: LLaVA became the most influential open-source VLM architecture, spawning LLaVA-1.5, LLaVA-NeXT, and countless derivatives through 2024.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack