DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Feng, Hao, Liu, Qi, Liu, Hao, Tang, Jingqun, Zhou, Wengang, Li, Houqiang, Huang, Can
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917846590685184
author Feng, Hao
Liu, Qi
Liu, Hao
Tang, Jingqun
Zhou, Wengang
Li, Houqiang
Huang, Can
author_facet Feng, Hao
Liu, Qi
Liu, Hao
Tang, Jingqun
Zhou, Wengang
Li, Houqiang
Huang, Can
contents This work presents DocPedia, a novel large multimodal model (LMM) for versatile OCR-free document understanding, capable of parsing images up to 2,560$\times$2,560 resolution. Unlike existing work either struggle with high-resolution documents or give up the large language model thus vision or language ability constrained, our DocPedia directly processes visual input in the frequency domain rather than the pixel space. The unique characteristic enables DocPedia to capture a greater amount of visual and textual information using a limited number of visual tokens. To consistently enhance both perception and comprehension abilities of our model, we develop a dual-stage training strategy and enrich instructions/annotations of all training tasks covering multiple document types. Extensive quantitative and qualitative experiments conducted on various publicly available benchmarks confirm the mutual benefits of jointly learning perception and comprehension tasks. The results provide further evidence of the effectiveness and superior performance of our DocPedia over other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2311_11810
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
Feng, Hao
Liu, Qi
Liu, Hao
Tang, Jingqun
Zhou, Wengang
Li, Houqiang
Huang, Can
Computer Vision and Pattern Recognition
Artificial Intelligence
This work presents DocPedia, a novel large multimodal model (LMM) for versatile OCR-free document understanding, capable of parsing images up to 2,560$\times$2,560 resolution. Unlike existing work either struggle with high-resolution documents or give up the large language model thus vision or language ability constrained, our DocPedia directly processes visual input in the frequency domain rather than the pixel space. The unique characteristic enables DocPedia to capture a greater amount of visual and textual information using a limited number of visual tokens. To consistently enhance both perception and comprehension abilities of our model, we develop a dual-stage training strategy and enrich instructions/annotations of all training tasks covering multiple document types. Extensive quantitative and qualitative experiments conducted on various publicly available benchmarks confirm the mutual benefits of jointly learning perception and comprehension tasks. The results provide further evidence of the effectiveness and superior performance of our DocPedia over other methods.
title DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2311.11810