Qwen-VL
Free while signed in. Answers cite the passages they came from.

Alibaba's Qwen-VL is a large-scale vision-language model family with strong performance across captioning, VQA, and visual localization.
Broad capability: Handles image captioning, visual QA, visual localization (grounding), and flexible multi-turn visual interaction.
Multilingual VL: Strong in both Chinese and English for visual tasks, filling a multilingual gap in VLMs predominantly English at the time.
Visual grounding: Supports bounding-box output for visual grounding, a capability not universally present in early VLMs.
Open release: Released as open weights, providing a strong open VLM baseline and kicking off the Qwen-VL family that has continued through 2024.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack