KOSMOS-2.5
Free while signed in. Answers cite the passages they came from.

Microsoft's KOSMOS-2.5 is a multimodal model purpose-built for "machine reading" of text-intensive images.
Text-rich image input: Specialized for documents, forms, receipts, and other images dominated by text rather than natural-scene imagery.
Document-level generation: Capable of document-level text generation from images, handling layout-aware reading order and structure.
Image-to-markdown: Converts complex text-rich images directly into Markdown output, preserving headings, lists, and tables.
Complements KOSMOS-1/2: Extends the KOSMOS family toward document intelligence, a domain where general VLMs had weaker performance.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack