ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sibue, Mathieu, Garza, Andres Muñoz, Mensah, Samuel, Shetty, Pranav, Ma, Zhiqiang, Liu, Xiaomo, Veloso, Manuela
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912900918018048
author Sibue, Mathieu
Garza, Andres Muñoz
Mensah, Samuel
Shetty, Pranav
Ma, Zhiqiang
Liu, Xiaomo
Veloso, Manuela
author_facet Sibue, Mathieu
Garza, Andres Muñoz
Mensah, Samuel
Shetty, Pranav
Ma, Zhiqiang
Liu, Xiaomo
Veloso, Manuela
contents Enterprise documents, such as forms and reports, embed critical information for downstream applications like data archiving, automated workflows, and analytics. Although generalist Vision Language Models (VLMs) perform well on established document understanding benchmarks, their ability to conduct holistic, fine-grained structured extraction across diverse document types and flexible schemas is not well studied. Existing Key Entity Extraction (KEE), Relation Extraction (RE), and Visual Question Answering (VQA) datasets are limited by narrow entity ontologies, simple queries, or homogeneous document types, often overlooking the need for adaptable and structured extraction. To address these gaps, we introduce ExStrucTiny, a new benchmark dataset for structured Information Extraction (IE) from document images, unifying aspects of KEE, RE, and VQA. Built through a novel pipeline combining manual and synthetic human-validated samples, ExStrucTiny covers more varied document types and extraction scenarios. We analyze open and closed VLMs on this benchmark, highlighting challenges such as schema adaptation, query under-specification, and answer localization. We hope our work provides a bedrock for improving generalist models for structured IE in documents.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12203
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
Sibue, Mathieu
Garza, Andres Muñoz
Mensah, Samuel
Shetty, Pranav
Ma, Zhiqiang
Liu, Xiaomo
Veloso, Manuela
Computation and Language
Enterprise documents, such as forms and reports, embed critical information for downstream applications like data archiving, automated workflows, and analytics. Although generalist Vision Language Models (VLMs) perform well on established document understanding benchmarks, their ability to conduct holistic, fine-grained structured extraction across diverse document types and flexible schemas is not well studied. Existing Key Entity Extraction (KEE), Relation Extraction (RE), and Visual Question Answering (VQA) datasets are limited by narrow entity ontologies, simple queries, or homogeneous document types, often overlooking the need for adaptable and structured extraction. To address these gaps, we introduce ExStrucTiny, a new benchmark dataset for structured Information Extraction (IE) from document images, unifying aspects of KEE, RE, and VQA. Built through a novel pipeline combining manual and synthetic human-validated samples, ExStrucTiny covers more varied document types and extraction scenarios. We analyze open and closed VLMs on this benchmark, highlighting challenges such as schema adaptation, query under-specification, and answer localization. We hope our work provides a bedrock for improving generalist models for structured IE in documents.
title ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
topic Computation and Language
url https://arxiv.org/abs/2602.12203