Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Jaeyoo, Choi, Jin Young, Park, Jeonghyung, Han, Bohyung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912110882062336
author Park, Jaeyoo
Choi, Jin Young
Park, Jeonghyung
Han, Bohyung
author_facet Park, Jaeyoo
Choi, Jin Young
Park, Jeonghyung
Han, Bohyung
contents We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale visual inputs for MLLMs, we propose the Hierarchical Visual Feature Aggregation (HVFA) module, designed to reduce the number of input tokens to LLMs. Leveraging a feature pyramid with cross-attentive pooling, our approach effectively manages the trade-off between information loss and efficiency without being affected by varying document image sizes. Furthermore, we introduce a novel instruction tuning task, which facilitates the model's text-reading capability by learning to predict the relative positions of input text, eventually minimizing the risk of truncated text caused by the limited capacity of LLMs. Comprehensive experiments validate the effectiveness of our approach, demonstrating superior performance in various document understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05254
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
Park, Jaeyoo
Choi, Jin Young
Park, Jeonghyung
Han, Bohyung
Computer Vision and Pattern Recognition
We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale visual inputs for MLLMs, we propose the Hierarchical Visual Feature Aggregation (HVFA) module, designed to reduce the number of input tokens to LLMs. Leveraging a feature pyramid with cross-attentive pooling, our approach effectively manages the trade-off between information loss and efficiency without being affected by varying document image sizes. Furthermore, we introduce a novel instruction tuning task, which facilitates the model's text-reading capability by learning to predict the relative positions of input text, eventually minimizing the risk of truncated text caused by the limited capacity of LLMs. Comprehensive experiments validate the effectiveness of our approach, demonstrating superior performance in various document understanding tasks.
title Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.05254