Visual Instruction Tuning (LLaVA)
Free while signed in. Answers cite the passages they came from.

Uses language-only GPT-4 to generate multimodal instruction-following data.
GPT-4-generated multimodal data: Bootstraps multimodal instruction data using only language-only GPT-4 given captions and bounding boxes - no direct visual access needed.
End-to-end training: Introduces LLaVA, an end-to-end trained large multimodal model combining CLIP vision encoder and Vicuna LLM.
Lightweight architecture: Simple projection layer between vision encoder and LLM - cheap and effective.
Open VLM revolution: LLaVA became the most influential open-source VLM architecture, spawning LLaVA-1.5, LLaVA-NeXT, and countless derivatives through 2024.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack