🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

DocLLM

Free while signed in. Answers cite the passages they came from.

First page
DocLLM
The curator’s take

JPMorgan's DocLLM is a lightweight extension to LLMs for visual-document understanding that uses bounding-box spatial information rather than image pixels.

Key points
01

Bounding-box attention: Incorporates 2D spatial layout by extending self-attention to condition on text bounding boxes, without needing an image encoder.

02

Irregular-layout pretraining: A custom pretraining objective handles the messy, heterogeneous layouts of real-world documents - forms, invoices, contracts - far better than naive text-only baselines.

03

Instruction tuning: The pretrained model is instruction-tuned on a document-intelligence dataset covering multiple document tasks.

04

SoTA on 14 of 16: Reports state-of-the-art performance on 14 out of 16 document-intelligence benchmarks, covering extraction, classification, layout understanding, and document QA.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack