InstructBLIP
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsVisual-language instruction tuning built on BLIP-2.
01
Instruction-aware Q-Former: Extends BLIP-2's Q-Former to be instruction-aware, dynamically extracting relevant visual features per instruction.
02
13 held-out datasets: Achieves state-of-the-art zero-shot performance on 13 held-out vision-language datasets.
03
Beats BLIP-2 and Flamingo: Outperforms both BLIP-2 and Flamingo on most zero-shot benchmarks despite being a direct BLIP-2 extension.
04
Open VLM progress: A prominent open-source VLM in 2023 that informed the later LLaVA-1.5, Qwen-VL, and InternVL lineage.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack