Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Ao, Gu, Zebo, Sun, Tenghao, Chen, Jiawen, Tu, Mingsheng, Cheng, Zifeng, Yin, Yafeng, Jiang, Zhiwei, Gu, Qing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914000888922112
author Zhou, Ao
Gu, Zebo
Sun, Tenghao
Chen, Jiawen
Tu, Mingsheng
Cheng, Zifeng
Yin, Yafeng
Jiang, Zhiwei
Gu, Qing
author_facet Zhou, Ao
Gu, Zebo
Sun, Tenghao
Chen, Jiawen
Tu, Mingsheng
Cheng, Zifeng
Yin, Yafeng
Jiang, Zhiwei
Gu, Qing
contents Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding capabilities in Visual Question Answering (VQA) tasks by integrating visual and textual features. However, under the challenging ten-choice question evaluation paradigm, existing methods still exhibit significant limitations when processing PDF documents with complex layouts and lengthy content. Notably, current mainstream models suffer from a strong bias toward English training data, resulting in suboptimal performance for Japanese and other language scenarios. To address these challenges, this paper proposes a novel Japanese PDF document understanding framework that combines multimodal hierarchical reasoning mechanisms with Colqwen-optimized retrieval methods, while innovatively introducing a semantic verification strategy through sub-question decomposition. Experimental results demonstrate that our framework not only significantly enhances the model's deep semantic parsing capability for complex documents, but also exhibits superior robustness in practical application scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16148
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
Zhou, Ao
Gu, Zebo
Sun, Tenghao
Chen, Jiawen
Tu, Mingsheng
Cheng, Zifeng
Yin, Yafeng
Jiang, Zhiwei
Gu, Qing
Information Retrieval
Computation and Language
Multimedia
Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding capabilities in Visual Question Answering (VQA) tasks by integrating visual and textual features. However, under the challenging ten-choice question evaluation paradigm, existing methods still exhibit significant limitations when processing PDF documents with complex layouts and lengthy content. Notably, current mainstream models suffer from a strong bias toward English training data, resulting in suboptimal performance for Japanese and other language scenarios. To address these challenges, this paper proposes a novel Japanese PDF document understanding framework that combines multimodal hierarchical reasoning mechanisms with Colqwen-optimized retrieval methods, while innovatively introducing a semantic verification strategy through sub-question decomposition. Experimental results demonstrate that our framework not only significantly enhances the model's deep semantic parsing capability for complex documents, but also exhibits superior robustness in practical application scenarios.
title Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
topic Information Retrieval
Computation and Language
Multimedia
url https://arxiv.org/abs/2508.16148