DocLLM
Free while signed in. Answers cite the passages they came from.

JPMorgan's DocLLM is a lightweight extension to LLMs for visual-document understanding that uses bounding-box spatial information rather than image pixels.
Bounding-box attention: Incorporates 2D spatial layout by extending self-attention to condition on text bounding boxes, without needing an image encoder.
Irregular-layout pretraining: A custom pretraining objective handles the messy, heterogeneous layouts of real-world documents - forms, invoices, contracts - far better than naive text-only baselines.
Instruction tuning: The pretrained model is instruction-tuned on a document-intelligence dataset covering multiple document tasks.
SoTA on 14 of 16: Reports state-of-the-art performance on 14 out of 16 document-intelligence benchmarks, covering extraction, classification, layout understanding, and document QA.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack