Enhancing Document Key Information Localization Through Data Augmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Dai, Yue
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915144544550912
author Dai, Yue
author_facet Dai, Yue
contents The Visually Rich Form Document Intelligence and Understanding (VRDIU) Track B focuses on the localization of key information in document images. The goal is to develop a method capable of localizing objects in both digital and handwritten documents, using only digital documents for training. This paper presents a simple yet effective approach that includes a document augmentation phase and an object detection phase. Specifically, we augment the training set of digital documents by mimicking the appearance of handwritten documents. Our experiments demonstrate that this pipeline enhances the models' generalization ability and achieves high performance in the competition.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06132
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Document Key Information Localization Through Data Augmentation
Dai, Yue
Computer Vision and Pattern Recognition
Computation and Language
The Visually Rich Form Document Intelligence and Understanding (VRDIU) Track B focuses on the localization of key information in document images. The goal is to develop a method capable of localizing objects in both digital and handwritten documents, using only digital documents for training. This paper presents a simple yet effective approach that includes a document augmentation phase and an object detection phase. Specifically, we augment the training set of digital documents by mimicking the appearance of handwritten documents. Our experiments demonstrate that this pipeline enhances the models' generalization ability and achieves high performance in the competition.
title Enhancing Document Key Information Localization Through Data Augmentation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2502.06132