Visually Guided Generative Text-Layout Pre-training for Document Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Zhiming, Bai, Haoli, Hou, Lu, Wei, Jiansheng, Jiang, Xin, Liu, Qun, Wong, Kam-Fai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913285968756736
author Mao, Zhiming
Bai, Haoli
Hou, Lu
Wei, Jiansheng
Jiang, Xin
Liu, Qun
Wong, Kam-Fai
author_facet Mao, Zhiming
Bai, Haoli
Hou, Lu
Wei, Jiansheng
Jiang, Xin
Liu, Qun
Wong, Kam-Fai
contents Prior study shows that pre-training techniques can boost the performance of visual document understanding (VDU), which typically requires models to gain abilities to perceive and reason both document texts and layouts (e.g., locations of texts and table-cells). To this end, we propose visually guided generative text-layout pre-training, named ViTLP. Given a document image, the model optimizes hierarchical language and layout modeling objectives to generate the interleaved text and layout sequence. In addition, to address the limitation of processing long documents by Transformers, we introduce a straightforward yet effective multi-segment generative pre-training scheme, facilitating ViTLP to process word-intensive documents of any length. ViTLP can function as a native OCR model to localize and recognize texts of document images. Besides, ViTLP can be effectively applied to various downstream VDU tasks. Extensive experiments show that ViTLP achieves competitive performance over existing baselines on benchmark VDU tasks, including information extraction, document classification, and document question answering.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16516
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visually Guided Generative Text-Layout Pre-training for Document Intelligence
Mao, Zhiming
Bai, Haoli
Hou, Lu
Wei, Jiansheng
Jiang, Xin
Liu, Qun
Wong, Kam-Fai
Computation and Language
Computer Vision and Pattern Recognition
Prior study shows that pre-training techniques can boost the performance of visual document understanding (VDU), which typically requires models to gain abilities to perceive and reason both document texts and layouts (e.g., locations of texts and table-cells). To this end, we propose visually guided generative text-layout pre-training, named ViTLP. Given a document image, the model optimizes hierarchical language and layout modeling objectives to generate the interleaved text and layout sequence. In addition, to address the limitation of processing long documents by Transformers, we introduce a straightforward yet effective multi-segment generative pre-training scheme, facilitating ViTLP to process word-intensive documents of any length. ViTLP can function as a native OCR model to localize and recognize texts of document images. Besides, ViTLP can be effectively applied to various downstream VDU tasks. Extensive experiments show that ViTLP achieves competitive performance over existing baselines on benchmark VDU tasks, including information extraction, document classification, and document question answering.
title Visually Guided Generative Text-Layout Pre-training for Document Intelligence
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.16516