🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

KOSMOS-2.5

First page
KOSMOS-2.5
Paper summary

Microsoft's KOSMOS-2.5 is a multimodal model purpose-built for "machine reading" of text-intensive images.

Ask this paper

Key points
01

Text-rich image input: Specialized for documents, forms, receipts, and other images dominated by text rather than natural-scene imagery.

02

Document-level generation: Capable of document-level text generation from images, handling layout-aware reading order and structure.

03

Image-to-markdown: Converts complex text-rich images directly into Markdown output, preserving headings, lists, and tables.

04

Complements KOSMOS-1/2: Extends the KOSMOS family toward document intelligence, a domain where general VLMs had weaker performance.

Every Monday
Get next week’s papers.
Subscribe on Substack