When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Tuo, Hu, Zhe, Li, Jing, Zhang, Hao, Lu, Yiren, Zhou, Yunlai, Qiao, Yiran, Liu, Disheng, Peng, Jeirui, Ma, Jing, Yin, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914472946302976
author Liang, Tuo
Hu, Zhe
Li, Jing
Zhang, Hao
Lu, Yiren
Zhou, Yunlai
Qiao, Yiran
Liu, Disheng
Peng, Jeirui
Ma, Jing
Yin, Yu
author_facet Liang, Tuo
Hu, Zhe
Li, Jing
Zhang, Hao
Lu, Yiren
Zhou, Yunlai
Qiao, Yiran
Liu, Disheng
Peng, Jeirui
Ma, Jing
Yin, Yu
contents Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant challenge for large vision-language models (VLMs). This limitation hinders AI's ability to engage in human-like reasoning and cultural expression. In this paper, we investigate this challenge through an in-depth analysis of comics that juxtapose panels to create humor through contradictions. We introduce the YesBut (V2), a novel benchmark with 1,262 comic images from diverse multilingual and multicultural contexts, featuring comprehensive annotations that capture various aspects of narrative understanding. Using this benchmark, we systematically evaluate a wide range of VLMs through four complementary tasks spanning from surface content comprehension to deep narrative reasoning, with particular emphasis on comparative reasoning between contradictory elements. Our extensive experiments reveal that even the most advanced models significantly underperform compared to humans, with common failures in visual perception, key element identification, comparative analysis and hallucinations. We further investigate text-based training strategies and social knowledge augmentation methods to enhance model performance. Our findings not only highlight critical weaknesses in VLMs' understanding of cultural and creative expressions but also provide pathways toward developing context-aware models capable of deeper narrative understanding though comparative reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
Liang, Tuo
Hu, Zhe
Li, Jing
Zhang, Hao
Lu, Yiren
Zhou, Yunlai
Qiao, Yiran
Liu, Disheng
Peng, Jeirui
Ma, Jing
Yin, Yu
Computer Vision and Pattern Recognition
Computation and Language
Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant challenge for large vision-language models (VLMs). This limitation hinders AI's ability to engage in human-like reasoning and cultural expression. In this paper, we investigate this challenge through an in-depth analysis of comics that juxtapose panels to create humor through contradictions. We introduce the YesBut (V2), a novel benchmark with 1,262 comic images from diverse multilingual and multicultural contexts, featuring comprehensive annotations that capture various aspects of narrative understanding. Using this benchmark, we systematically evaluate a wide range of VLMs through four complementary tasks spanning from surface content comprehension to deep narrative reasoning, with particular emphasis on comparative reasoning between contradictory elements. Our extensive experiments reveal that even the most advanced models significantly underperform compared to humans, with common failures in visual perception, key element identification, comparative analysis and hallucinations. We further investigate text-based training strategies and social knowledge augmentation methods to enhance model performance. Our findings not only highlight critical weaknesses in VLMs' understanding of cultural and creative expressions but also provide pathways toward developing context-aware models capable of deeper narrative understanding though comparative reasoning.
title When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.23137