The Dawn of LMMs (GPT-4V Deep Dive)
Free while signed in. Answers cite the passages they came from.

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.
Comprehensive task coverage: Probes GPT-4V across visual reasoning, code, OCR, document understanding, multimodal commonsense, and agent-style tasks.
Working input modes: Catalogs the diverse input patterns GPT-4V supports - single images, multi-image reasoning, image-text interleaving, sketches, and handwritten input.
Capability frontier: Demonstrates emergent capabilities like reading diagrams, interpreting medical imaging, and extracting structured information from complex visuals.
Open issues: Identifies persistent weaknesses including hallucination, fine-grained spatial reasoning, and consistency across related queries - a reference for what was still broken at the start of the GPT-4V era.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack