GraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Panfeng, Wu, Jian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914962466668544
author Cao, Panfeng
Wu, Jian
author_facet Cao, Panfeng
Wu, Jian
contents Key information extraction (KIE) from visually rich documents (VRD) has been a challenging task in document intelligence because of not only the complicated and diverse layouts of VRD that make the model hard to generalize but also the lack of methods to exploit the multimodal features in VRD. In this paper, we propose a light-weight model named GraphRevisedIE that effectively embeds multimodal features such as textual, visual, and layout features from VRD and leverages graph revision and graph convolution to enrich the multimodal embedding with global context. Extensive experiments on multiple real-world datasets show that GraphRevisedIE generalizes to documents of varied layouts and achieves comparable or better performance compared to previous KIE methods. We also publish a business license dataset that contains both real-life and synthesized documents to facilitate research of document KIE.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network
Cao, Panfeng
Wu, Jian
Information Retrieval
Computer Vision and Pattern Recognition
Key information extraction (KIE) from visually rich documents (VRD) has been a challenging task in document intelligence because of not only the complicated and diverse layouts of VRD that make the model hard to generalize but also the lack of methods to exploit the multimodal features in VRD. In this paper, we propose a light-weight model named GraphRevisedIE that effectively embeds multimodal features such as textual, visual, and layout features from VRD and leverages graph revision and graph convolution to enrich the multimodal embedding with global context. Extensive experiments on multiple real-world datasets show that GraphRevisedIE generalizes to documents of varied layouts and achieves comparable or better performance compared to previous KIE methods. We also publish a business license dataset that contains both real-life and synthesized documents to facilitate research of document KIE.
title GraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network
topic Information Retrieval
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.01160