8/ DocLLM – an extension to traditional LLMs for reasoning over visual documents; focuses on using bounding box information to incorporate spatial layout structure; demonstrates SoTA on 14 of 16 datasets across several document intelligence tasks.
DocLLM: Visual Document Reasoning with Bounding Box Spatial Layout
By
–