🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

The Dawn of LMMs (GPT-4V Deep Dive)

First page
The Dawn of LMMs (GPT-4V Deep Dive)
Paper summary

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.

Ask this paper

Key points
01

Comprehensive task coverage: Probes GPT-4V across visual reasoning, code, OCR, document understanding, multimodal commonsense, and agent-style tasks.

02

Working input modes: Catalogs the diverse input patterns GPT-4V supports - single images, multi-image reasoning, image-text interleaving, sketches, and handwritten input.

03

Capability frontier: Demonstrates emergent capabilities like reading diagrams, interpreting medical imaging, and extracting structured information from complex visuals.

04

Open issues: Identifies persistent weaknesses including hallucination, fine-grained spatial reasoning, and consistency across related queries - a reference for what was still broken at the start of the GPT-4V era.

Every Monday
Get next week’s papers.
Subscribe on Substack