🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Training

InstructBLIP

First page
InstructBLIP
Paper summary

Visual-language instruction tuning built on BLIP-2.

Ask this paper

Key points
01

Instruction-aware Q-Former: Extends BLIP-2's Q-Former to be instruction-aware, dynamically extracting relevant visual features per instruction.

02

13 held-out datasets: Achieves state-of-the-art zero-shot performance on 13 held-out vision-language datasets.

03

Beats BLIP-2 and Flamingo: Outperforms both BLIP-2 and Flamingo on most zero-shot benchmarks despite being a direct BLIP-2 extension.

04

Open VLM progress: A prominent open-source VLM in 2023 that informed the later LLaVA-1.5, Qwen-VL, and InternVL lineage.

Every Monday
Get next week’s papers.
Subscribe on Substack