MultiModal-GPT
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsA vision-language model for multi-round dialogue fine-tuned from OpenFlamingo.
01
LoRA-based extension: Adds LoRA to OpenFlamingo's cross-attention and self-attention for efficient fine-tuning.
02
Multi-round dialog: Specifically designed for multi-turn visual dialog, going beyond single-turn VQA.
03
Open visual chatbot: An early fully-open visual chatbot that users could run locally.
04
VLM dialog research: Informed the trajectory toward modern visual chatbots (LLaVA, Qwen-VL) that dominated open VLM research.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack