DISCO: Document Intelligence Suite for COmparative Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benkirane, Kenza, Goldwater, Dan, Asenov, Martin, Ghodsi, Aneiss
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914420086538240
author Benkirane, Kenza
Goldwater, Dan
Asenov, Martin
Ghodsi, Aneiss
author_facet Benkirane, Kenza
Goldwater, Dan
Asenov, Martin
Ghodsi, Aneiss
contents Document intelligence requires accurate text extraction and reliable reasoning over document content. We introduce \textbf{DISCO}, a \emph{Document Intelligence Suite for COmparative Evaluation}, that evaluates optical character recognition (OCR) pipelines and vision-language models (VLMs) separately on parsing and question answering across diverse document types, including handwritten text, multilingual scripts, medical forms, infographics, and multi-page documents. Our evaluation shows that performance varies substantially across tasks and document characteristics, underscoring the need for complexity-aware approach selection. OCR pipelines are generally more reliable for handwriting and for long or multi-page documents, where explicit text grounding supports text-heavy reasoning, while VLMs perform better on multilingual text and visually rich layouts. Task-aware prompting yields mixed effects, improving performance on some document types while degrading it on others. These findings provide empirical guidance for selecting document processing strategies based on document structure and reasoning demands.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23511
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DISCO: Document Intelligence Suite for COmparative Evaluation
Benkirane, Kenza
Goldwater, Dan
Asenov, Martin
Ghodsi, Aneiss
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.7; H.3.3
Document intelligence requires accurate text extraction and reliable reasoning over document content. We introduce \textbf{DISCO}, a \emph{Document Intelligence Suite for COmparative Evaluation}, that evaluates optical character recognition (OCR) pipelines and vision-language models (VLMs) separately on parsing and question answering across diverse document types, including handwritten text, multilingual scripts, medical forms, infographics, and multi-page documents. Our evaluation shows that performance varies substantially across tasks and document characteristics, underscoring the need for complexity-aware approach selection. OCR pipelines are generally more reliable for handwriting and for long or multi-page documents, where explicit text grounding supports text-heavy reasoning, while VLMs perform better on multilingual text and visually rich layouts. Task-aware prompting yields mixed effects, improving performance on some document types while degrading it on others. These findings provide empirical guidance for selecting document processing strategies based on document structure and reasoning demands.
title DISCO: Document Intelligence Suite for COmparative Evaluation
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.7; H.3.3
url https://arxiv.org/abs/2603.23511