Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zining, Guan, Tongkun, Fu, Pei, Duan, Chen, Jiang, Qianyi, Guo, Zhentao, Guo, Shan, Luo, Junfeng, Shen, Wei, Yang, Xiaokang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916655975628800
author Wang, Zining
Guan, Tongkun
Fu, Pei
Duan, Chen
Jiang, Qianyi
Guo, Zhentao
Guo, Shan
Luo, Junfeng
Shen, Wei
Yang, Xiaokang
author_facet Wang, Zining
Guan, Tongkun
Fu, Pei
Duan, Chen
Jiang, Qianyi
Guo, Zhentao
Guo, Shan
Luo, Junfeng
Shen, Wei
Yang, Xiaokang
contents Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in document-level MLLMs remains underexplored. In this study, we introduce a novel visual-language alignment method that casts the key issue as a Visual Question Answering with Mask generation (VQAMask) task, optimizing two tasks simultaneously: VQA-based text parsing and mask generation. The former allows the model to implicitly align images and text at the semantic level. The latter introduces an additional mask generator (discarded during inference) to explicitly ensure alignment between visual texts within images and their corresponding image regions at a spatially-aware level. Together, they can prevent model hallucinations when parsing visual text and effectively promote spatially-aware feature representation learning. To support the proposed VQAMask task, we construct a comprehensive image-mask generation pipeline and provide a large-scale dataset with 6M data (MTMask6M). Subsequently, we demonstrate that introducing the proposed mask generation task yields competitive document-level understanding performance. Leveraging the proposed VQAMask, we introduce Marten, a training-efficient MLLM tailored for document-level understanding. Extensive experiments show that our Marten consistently achieves significant improvements among 8B-MLLMs in document-centric tasks. Code and datasets are available at https://github.com/PriNing/Marten.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14140
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
Wang, Zining
Guan, Tongkun
Fu, Pei
Duan, Chen
Jiang, Qianyi
Guo, Zhentao
Guo, Shan
Luo, Junfeng
Shen, Wei
Yang, Xiaokang
Computer Vision and Pattern Recognition
Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in document-level MLLMs remains underexplored. In this study, we introduce a novel visual-language alignment method that casts the key issue as a Visual Question Answering with Mask generation (VQAMask) task, optimizing two tasks simultaneously: VQA-based text parsing and mask generation. The former allows the model to implicitly align images and text at the semantic level. The latter introduces an additional mask generator (discarded during inference) to explicitly ensure alignment between visual texts within images and their corresponding image regions at a spatially-aware level. Together, they can prevent model hallucinations when parsing visual text and effectively promote spatially-aware feature representation learning. To support the proposed VQAMask task, we construct a comprehensive image-mask generation pipeline and provide a large-scale dataset with 6M data (MTMask6M). Subsequently, we demonstrate that introducing the proposed mask generation task yields competitive document-level understanding performance. Leveraging the proposed VQAMask, we introduce Marten, a training-efficient MLLM tailored for document-level understanding. Extensive experiments show that our Marten consistently achieves significant improvements among 8B-MLLMs in document-centric tasks. Code and datasets are available at https://github.com/PriNing/Marten.
title Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.14140