MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Baek, Jeonghun, Egashira, Kazuki, Onohara, Shota, Miyai, Atsuyuki, Imajuku, Yuki, Ikuta, Hikaru, Aizawa, Kiyoharu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2026)
by: Baek, Jeonghun, et al.
Published: (2026)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
by: Onohara, Shota, et al.
Published: (2024)
by: Onohara, Shota, et al.
Published: (2024)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
by: Ikuta, Hikaru, et al.
Published: (2024)
by: Ikuta, Hikaru, et al.
Published: (2024)
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
by: Li, Yingxuan, et al.
Published: (2023)
by: Li, Yingxuan, et al.
Published: (2023)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
Context-Informed Machine Translation of Manga using Multimodal Large Language Models
by: Lippmann, Philip, et al.
Published: (2024)
by: Lippmann, Philip, et al.
Published: (2024)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
by: Kawakami, Tatsuki, et al.
Published: (2025)
by: Kawakami, Tatsuki, et al.
Published: (2025)
Re:Verse -- Can Your VLM Read a Manga?
by: Baranwal, Aaditya, et al.
Published: (2025)
by: Baranwal, Aaditya, et al.
Published: (2025)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
by: Imajuku, Yuki, et al.
Published: (2024)
by: Imajuku, Yuki, et al.
Published: (2024)
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
by: Wang, Muyao, et al.
Published: (2026)
by: Wang, Muyao, et al.
Published: (2026)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
Memes-as-Replies: Can Models Select Humorous Manga Panel Responses?
by: Kohita, Ryosuke, et al.
Published: (2026)
by: Kohita, Ryosuke, et al.
Published: (2026)
You Only Look at Manga Faces: a YOLOv12-based Approach on Manga Datasets and Benchmark
by: Fargetta, Georgia, et al.
Published: (2025)
by: Fargetta, Georgia, et al.
Published: (2025)
Sketch2Manga: Shaded Manga Screening from Sketch with Diffusion Models
by: Lin, Jian, et al.
Published: (2024)
by: Lin, Jian, et al.
Published: (2024)
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
by: Miyai, Atsuyuki, et al.
Published: (2026)
by: Miyai, Atsuyuki, et al.
Published: (2026)
GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
by: Miyai, Atsuyuki, et al.
Published: (2023)
by: Miyai, Atsuyuki, et al.
Published: (2023)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
by: Miyai, Atsuyuki, et al.
Published: (2024)
by: Miyai, Atsuyuki, et al.
Published: (2024)
A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
by: Noda, Shiho, et al.
Published: (2025)
by: Noda, Shiho, et al.
Published: (2025)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
Manga Generation via Layout-controllable Diffusion
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Journal of Anime and Manga Studies
Published: (2022)
Published: (2022)
Manga Vision
Published: (2018)
Published: (2018)
Inference-time Trajectory Optimization for Manga Image Editing
by: Furuta, Ryosuke
Published: (2026)
by: Furuta, Ryosuke
Published: (2026)
Number it: Temporal Grounding Videos like Flipping Manga
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
MangaNinja: Line Art Colorization with Precise Reference Following
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
A Manga Perfeita
by: Manning, Erin
Published: (2019)
by: Manning, Erin
Published: (2019)
Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
Region-Wise Correspondence Prediction between Manga Line Art Images
by: Li, Yingxuan, et al.
Published: (2025)
by: Li, Yingxuan, et al.
Published: (2025)
Selecting Mangas and Graphic Novels
by: Nylund, Carol
Published: (2007)
by: Nylund, Carol
Published: (2007)
How Panel Layouts Define Manga: Insights from Visual Ablation Experiments
by: Feng, Siyuan, et al.
Published: (2024)
by: Feng, Siyuan, et al.
Published: (2024)
GANime: Generating Anime and Manga Character Drawings from Sketches with Deep Learning
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers
by: Qiu, Qianru, et al.
Published: (2025)
by: Qiu, Qianru, et al.
Published: (2025)
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
by: Taniguchi, Takara, et al.
Published: (2024)
by: Taniguchi, Takara, et al.
Published: (2024)
Panel-by-Panel Souls: A Performative Workflow for Expressive Faces in AI-Assisted Manga Creation
by: Zhang, Qing, et al.
Published: (2025)
by: Zhang, Qing, et al.
Published: (2025)
A Layman's Lexicon of Manga and Anime
by: Bunche, Steve
Published: (2004)
by: Bunche, Steve
Published: (2004)
Journey into the World of Manga and Graphic Novels
by: Simmons, Mary, et al.
Published: (2009)
by: Simmons, Mary, et al.
Published: (2009)
Sportoonizer: Augmenting Sports Highlights' Narration and Visual Impact via Automatic Manga B-Roll Generation
by: Hu, Siying, et al.
Published: (2024)
by: Hu, Siying, et al.
Published: (2024)
Similar Items
-
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2026) -
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
by: Onohara, Shota, et al.
Published: (2024) -
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
by: Ikuta, Hikaru, et al.
Published: (2024) -
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
by: Miyai, Atsuyuki, et al.
Published: (2025) -
Harnessing PDF Data for Improving Japanese Large Multimodal Models
by: Baek, Jeonghun, et al.
Published: (2025)