v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Zhengpeng, Zhao, Yanpeng, Zhou, Jianqun, Wang, Yuxuan, Cui, Qinrong, Bi, Wei, Zhu, Songchun, Zhao, Bo, Zheng, Zilong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914622415568896
author Shi, Zhengpeng
Zhao, Yanpeng
Zhou, Jianqun
Wang, Yuxuan
Cui, Qinrong
Bi, Wei
Zhu, Songchun
Zhao, Bo
Zheng, Zilong
author_facet Shi, Zhengpeng
Zhao, Yanpeng
Zhou, Jianqun
Wang, Yuxuan
Cui, Qinrong
Bi, Wei
Zhu, Songchun
Zhao, Bo
Zheng, Zilong
contents AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimodal large language models (MLLMs) for humor understanding, we introduce v-HUB, a novel video humor understanding benchmark. v-HUB comprises a curated collection of non-verbal short videos, reflecting real-world scenarios where humor can be appreciated purely through visual cues. We pair each video clip with rich annotations to support a variety of evaluation tasks and analyses, including a novel study of environmental sound that can enhance humor. To broaden its applicability, we construct an open-ended QA task, making v-HUB readily integrable into existing video understanding task suites. We evaluate a diverse set of MLLMs, from specialized Video-LLMs to versatile OmniLLMs that can natively process audio, covering both open-source and proprietary domains. The experimental results expose the difficulties MLLMs face in comprehending humor from visual cues alone. Our findings also demonstrate that incorporating audio helps with video humor understanding, highlighting the promise of integrating richer modalities for complex video understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
Shi, Zhengpeng
Zhao, Yanpeng
Zhou, Jianqun
Wang, Yuxuan
Cui, Qinrong
Bi, Wei
Zhu, Songchun
Zhao, Bo
Zheng, Zilong
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimodal large language models (MLLMs) for humor understanding, we introduce v-HUB, a novel video humor understanding benchmark. v-HUB comprises a curated collection of non-verbal short videos, reflecting real-world scenarios where humor can be appreciated purely through visual cues. We pair each video clip with rich annotations to support a variety of evaluation tasks and analyses, including a novel study of environmental sound that can enhance humor. To broaden its applicability, we construct an open-ended QA task, making v-HUB readily integrable into existing video understanding task suites. We evaluate a diverse set of MLLMs, from specialized Video-LLMs to versatile OmniLLMs that can natively process audio, covering both open-source and proprietary domains. The experimental results expose the difficulties MLLMs face in comprehending humor from visual cues alone. Our findings also demonstrate that incorporating audio helps with video humor understanding, highlighting the promise of integrating richer modalities for complex video understanding tasks.
title v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.25773