🚀NEW LABGetting Started with Claude AgentsStart lab
Training

DocLLM

First page
DocLLM
Paper summary

JPMorgan's DocLLM is a lightweight extension to LLMs for visual-document understanding that uses bounding-box spatial information rather than image pixels.

Ask this paper

Key points
01

Bounding-box attention: Incorporates 2D spatial layout by extending self-attention to condition on text bounding boxes, without needing an image encoder.

02

Irregular-layout pretraining: A custom pretraining objective handles the messy, heterogeneous layouts of real-world documents - forms, invoices, contracts - far better than naive text-only baselines.

03

Instruction tuning: The pretrained model is instruction-tuned on a document-intelligence dataset covering multiple document tasks.

04

SoTA on 14 of 16: Reports state-of-the-art performance on 14 out of 16 document-intelligence benchmarks, covering extraction, classification, layout understanding, and document QA.

Every Monday
Get next week’s papers.
Subscribe on Substack