Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Chuwei, Tang, Guozhi, Zheng, Qi, Yao, Cong, Jin, Lianwen, Li, Chenliang, Xue, Yang, Si, Luo
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915349080834048
author Luo, Chuwei
Tang, Guozhi
Zheng, Qi
Yao, Cong
Jin, Lianwen
Li, Chenliang
Xue, Yang
Si, Luo
author_facet Luo, Chuwei
Tang, Guozhi
Zheng, Qi
Yao, Cong
Jin, Lianwen
Li, Chenliang
Xue, Yang
Si, Luo
contents Multi-modal document pre-trained models have proven to be very effective in a variety of visually-rich document understanding (VrDU) tasks. Though existing document pre-trained models have achieved excellent performance on standard benchmarks for VrDU, the way they model and exploit the interactions between vision and language on documents has hindered them from better generalization ability and higher accuracy. In this work, we investigate the problem of vision-language joint representation learning for VrDU mainly from the perspective of supervisory signals. Specifically, a pre-training paradigm called Bi-VLDoc is proposed, in which a bidirectional vision-language supervision strategy and a vision-language hybrid-attention mechanism are devised to fully explore and utilize the interactions between these two modalities, to learn stronger cross-modal document representations with richer semantics. Benefiting from the learned informative cross-modal document representations, Bi-VLDoc significantly advances the state-of-the-art performance on three widely-used document understanding benchmarks, including Form Understanding (from 85.14% to 93.44%), Receipt Information Extraction (from 96.01% to 97.84%), and Document Classification (from 96.08% to 97.12%). On Document Visual QA, Bi-VLDoc achieves the state-of-the-art performance compared to previous single model methods.
format Preprint
id arxiv_https___arxiv_org_abs_2206_13155
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
Luo, Chuwei
Tang, Guozhi
Zheng, Qi
Yao, Cong
Jin, Lianwen
Li, Chenliang
Xue, Yang
Si, Luo
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
Multi-modal document pre-trained models have proven to be very effective in a variety of visually-rich document understanding (VrDU) tasks. Though existing document pre-trained models have achieved excellent performance on standard benchmarks for VrDU, the way they model and exploit the interactions between vision and language on documents has hindered them from better generalization ability and higher accuracy. In this work, we investigate the problem of vision-language joint representation learning for VrDU mainly from the perspective of supervisory signals. Specifically, a pre-training paradigm called Bi-VLDoc is proposed, in which a bidirectional vision-language supervision strategy and a vision-language hybrid-attention mechanism are devised to fully explore and utilize the interactions between these two modalities, to learn stronger cross-modal document representations with richer semantics. Benefiting from the learned informative cross-modal document representations, Bi-VLDoc significantly advances the state-of-the-art performance on three widely-used document understanding benchmarks, including Form Understanding (from 85.14% to 93.44%), Receipt Information Extraction (from 96.01% to 97.84%), and Document Classification (from 96.08% to 97.12%). On Document Visual QA, Bi-VLDoc achieves the state-of-the-art performance compared to previous single model methods.
title Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2206.13155