MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baek, Jeonghun, Egashira, Kazuki, Onohara, Shota, Miyai, Atsuyuki, Imajuku, Yuki, Ikuta, Hikaru, Aizawa, Kiyoharu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912847343124480
author Baek, Jeonghun
Egashira, Kazuki
Onohara, Shota
Miyai, Atsuyuki
Imajuku, Yuki
Ikuta, Hikaru
Aizawa, Kiyoharu
author_facet Baek, Jeonghun
Egashira, Kazuki
Onohara, Shota
Miyai, Atsuyuki
Imajuku, Yuki
Ikuta, Hikaru
Aizawa, Kiyoharu
contents Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga creators reflect on and refine their stories. To this end, we introduce two benchmarks for multimodal manga understanding: MangaOCR, which targets in-page text recognition, and MangaVQA, a novel benchmark designed to evaluate contextual understanding through visual question answering. MangaVQA consists of 526 high-quality, manually constructed question-answer pairs, enabling reliable evaluation across diverse narrative and visual scenarios. Building on these benchmarks, we develop MangaLMM, a manga-specialized model finetuned from the open-source LMM Qwen2.5-VL to jointly handle both tasks. Through extensive experiments, including comparisons with proprietary models such as GPT-4o and Gemini 2.5, we assess how well LMMs understand manga. Our benchmark and model provide a comprehensive foundation for evaluating and advancing LMMs in the richly narrative domain of manga.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
Baek, Jeonghun
Egashira, Kazuki
Onohara, Shota
Miyai, Atsuyuki
Imajuku, Yuki
Ikuta, Hikaru
Aizawa, Kiyoharu
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga creators reflect on and refine their stories. To this end, we introduce two benchmarks for multimodal manga understanding: MangaOCR, which targets in-page text recognition, and MangaVQA, a novel benchmark designed to evaluate contextual understanding through visual question answering. MangaVQA consists of 526 high-quality, manually constructed question-answer pairs, enabling reliable evaluation across diverse narrative and visual scenarios. Building on these benchmarks, we develop MangaLMM, a manga-specialized model finetuned from the open-source LMM Qwen2.5-VL to jointly handle both tasks. Through extensive experiments, including comparisons with proprietary models such as GPT-4o and Gemini 2.5, we assess how well LMMs understand manga. Our benchmark and model provide a comprehensive foundation for evaluating and advancing LMMs in the richly narrative domain of manga.
title MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20298