DocLLM

JPMorgan's DocLLM is a lightweight extension to LLMs for visual-document understanding that uses bounding-box spatial information rather than image pixels.
Ask this paper
Bounding-box attention: Incorporates 2D spatial layout by extending self-attention to condition on text bounding boxes, without needing an image encoder.
Irregular-layout pretraining: A custom pretraining objective handles the messy, heterogeneous layouts of real-world documents - forms, invoices, contracts - far better than naive text-only baselines.
Instruction tuning: The pretrained model is instruction-tuned on a document-intelligence dataset covering multiple document tasks.
SoTA on 14 of 16: Reports state-of-the-art performance on 14 out of 16 document-intelligence benchmarks, covering extraction, classification, layout understanding, and document QA.