InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tanaka, Ryota, Iki, Taichi, Nishida, Kyosuke, Saito, Kuniko, Suzuki, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914650871824384
author Tanaka, Ryota
Iki, Taichi
Nishida, Kyosuke
Saito, Kuniko
Suzuki, Jun
author_facet Tanaka, Ryota
Iki, Taichi
Nishida, Kyosuke
Saito, Kuniko
Suzuki, Jun
contents We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions. To this end, we propose InstructDoc, the first large-scale collection of 30 publicly available VDU datasets, each with diverse instructions in a unified format, which covers a wide range of 12 tasks and includes open document types/formats. Furthermore, to enhance the generalization performance on VDU tasks, we design a new instruction-based document reading and understanding model, InstructDr, that connects document images, image encoders, and large language models (LLMs) through a trainable bridging module. Experiments demonstrate that InstructDr can effectively adapt to new VDU datasets, tasks, and domains via given instructions and outperforms existing multimodal LLMs and ChatGPT without specific training.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
Tanaka, Ryota
Iki, Taichi
Nishida, Kyosuke
Saito, Kuniko
Suzuki, Jun
Computer Vision and Pattern Recognition
Computation and Language
We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions. To this end, we propose InstructDoc, the first large-scale collection of 30 publicly available VDU datasets, each with diverse instructions in a unified format, which covers a wide range of 12 tasks and includes open document types/formats. Furthermore, to enhance the generalization performance on VDU tasks, we design a new instruction-based document reading and understanding model, InstructDr, that connects document images, image encoders, and large language models (LLMs) through a trainable bridging module. Experiments demonstrate that InstructDr can effectively adapt to new VDU datasets, tasks, and domains via given instructions and outperforms existing multimodal LLMs and ChatGPT without specific training.
title InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2401.13313