MangaUB: A Manga Understanding Benchmark for Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ikuta, Hikaru, Wöhler, Leslie, Aizawa, Kiyoharu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917811933151232
author Ikuta, Hikaru
Wöhler, Leslie
Aizawa, Kiyoharu
author_facet Ikuta, Hikaru
Wöhler, Leslie
Aizawa, Kiyoharu
contents Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panels differ from natural images, computational systems traditionally had to be designed specifically for manga. Recently, the adaptive nature of modern large multimodal models (LMMs) shows possibilities for more general approaches. To provide an analysis of the current capability of LMMs for manga understanding tasks and identifying areas for their improvement, we design and evaluate MangaUB, a novel manga understanding benchmark for LMMs. MangaUB is designed to assess the recognition and understanding of content shown in a single panel as well as conveyed across multiple panels, allowing for a fine-grained analysis of a model's various capabilities required for manga understanding. Our results show strong performance on the recognition of image content, while understanding the emotion and information conveyed across multiple panels is still challenging, highlighting future work towards LMMs for manga understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19034
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
Ikuta, Hikaru
Wöhler, Leslie
Aizawa, Kiyoharu
Computer Vision and Pattern Recognition
Multimedia
Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panels differ from natural images, computational systems traditionally had to be designed specifically for manga. Recently, the adaptive nature of modern large multimodal models (LMMs) shows possibilities for more general approaches. To provide an analysis of the current capability of LMMs for manga understanding tasks and identifying areas for their improvement, we design and evaluate MangaUB, a novel manga understanding benchmark for LMMs. MangaUB is designed to assess the recognition and understanding of content shown in a single panel as well as conveyed across multiple panels, allowing for a fine-grained analysis of a model's various capabilities required for manga understanding. Our results show strong performance on the recognition of image content, while understanding the emotion and information conveyed across multiple panels is still challenging, highlighting future work towards LMMs for manga understanding.
title MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2407.19034