VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Yihao, Han, Soyeon Caren, Li, Yan, Poon, Josiah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912409360269312
author Ding, Yihao
Han, Soyeon Caren
Li, Yan
Poon, Josiah
author_facet Ding, Yihao
Han, Soyeon Caren
Li, Yan
Poon, Josiah
contents Visually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across domains such as medical, financial, and educational applications. However, form-like documents pose unique challenges due to their complex layouts, multi-stakeholder involvement, and high structural variability. Addressing these issues, the VRD-IU Competition was introduced, focusing on extracting and localizing key information from multi-format forms within the Form-NLU dataset, which includes digital, printed, and handwritten documents. This paper presents insights from the competition, which featured two tracks: Track A, emphasizing entity-based key information retrieval, and Track B, targeting end-to-end key information localization from raw document images. With over 20 participating teams, the competition showcased various state-of-the-art methodologies, including hierarchical decomposition, transformer-based retrieval, multimodal feature fusion, and advanced object detection techniques. The top-performing models set new benchmarks in VRDU, providing valuable insights into document intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
Ding, Yihao
Han, Soyeon Caren
Li, Yan
Poon, Josiah
Computer Vision and Pattern Recognition
Artificial Intelligence
Visually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across domains such as medical, financial, and educational applications. However, form-like documents pose unique challenges due to their complex layouts, multi-stakeholder involvement, and high structural variability. Addressing these issues, the VRD-IU Competition was introduced, focusing on extracting and localizing key information from multi-format forms within the Form-NLU dataset, which includes digital, printed, and handwritten documents. This paper presents insights from the competition, which featured two tracks: Track A, emphasizing entity-based key information retrieval, and Track B, targeting end-to-end key information localization from raw document images. With over 20 participating teams, the competition showcased various state-of-the-art methodologies, including hierarchical decomposition, transformer-based retrieval, multimodal feature fusion, and advanced object detection techniques. The top-performing models set new benchmarks in VRDU, providing valuable insights into document intelligence.
title VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.01388