Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ryan, Yuriel, Tan, Rui Yang, Choo, Kenny Tsu Wei, Lee, Roy Ka-Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908543180865536
author Ryan, Yuriel
Tan, Rui Yang
Choo, Kenny Tsu Wei
Lee, Roy Ka-Wei
author_facet Ryan, Yuriel
Tan, Rui Yang
Choo, Kenny Tsu Wei
Lee, Roy Ka-Wei
contents Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs' ability to interpret multimodal humor and recognize narrative sequences. Experiments with state-of-the-art LMMs reveal substantial gaps: for instance, top models achieve only 61% accuracy in panel sequencing, far below human performance. This underscores critical limitations in current models' integration of visual and textual cues for coherent narrative and humor understanding. By providing a rigorous framework for evaluating multimodal contextual and narrative reasoning, PixelHumor aims to drive the development of LMMs that better engage in natural, socially aware interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
Ryan, Yuriel
Tan, Rui Yang
Choo, Kenny Tsu Wei
Lee, Roy Ka-Wei
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs' ability to interpret multimodal humor and recognize narrative sequences. Experiments with state-of-the-art LMMs reveal substantial gaps: for instance, top models achieve only 61% accuracy in panel sequencing, far below human performance. This underscores critical limitations in current models' integration of visual and textual cues for coherent narrative and humor understanding. By providing a rigorous framework for evaluating multimodal contextual and narrative reasoning, PixelHumor aims to drive the development of LMMs that better engage in natural, socially aware interactions.
title Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.12248