Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908543180865536 |
|---|---|
| author | Ryan, Yuriel Tan, Rui Yang Choo, Kenny Tsu Wei Lee, Roy Ka-Wei |
| author_facet | Ryan, Yuriel Tan, Rui Yang Choo, Kenny Tsu Wei Lee, Roy Ka-Wei |
| contents | Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs' ability to interpret multimodal humor and recognize narrative sequences. Experiments with state-of-the-art LMMs reveal substantial gaps: for instance, top models achieve only 61% accuracy in panel sequencing, far below human performance. This underscores critical limitations in current models' integration of visual and textual cues for coherent narrative and humor understanding. By providing a rigorous framework for evaluating multimodal contextual and narrative reasoning, PixelHumor aims to drive the development of LMMs that better engage in natural, socially aware interactions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_12248 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics Ryan, Yuriel Tan, Rui Yang Choo, Kenny Tsu Wei Lee, Roy Ka-Wei Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs' ability to interpret multimodal humor and recognize narrative sequences. Experiments with state-of-the-art LMMs reveal substantial gaps: for instance, top models achieve only 61% accuracy in panel sequencing, far below human performance. This underscores critical limitations in current models' integration of visual and textual cues for coherent narrative and humor understanding. By providing a rigorous framework for evaluating multimodal contextual and narrative reasoning, PixelHumor aims to drive the development of LMMs that better engage in natural, socially aware interactions. |
| title | Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2509.12248 |