Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Run, Li, Lu, Zhang, Rongzhao, Xu, Jie
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915945604186112
author Xu, Run
Li, Lu
Zhang, Rongzhao
Xu, Jie
author_facet Xu, Run
Li, Lu
Zhang, Rongzhao
Xu, Jie
contents Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over explicit cultural context, making it difficult to jointly maintain image relevance, contextual appropriateness, and humor quality under a specified cultural background. To address this limitation, we introduce a new multimodal generation task, culture-aware humorous captioning, which requires a model to generate a humorous caption conditioned on both an input image and a target cultural context. Captions generated under different cultural contexts are not expected to share the same surface form, but should remain grounded in similar visual situations or humorous rationales.To support this task, we establish a six-dimensional evaluation framework covering image relevance, contextual fit, semantic richness, reasonableness, humor, and creativity. We further propose a staged alignment framework that first initializes the model with high-resource supervision under the Western cultural context, then performs multi-dimensional preference alignment via judge-based GRPO with a Degradation-aware Prototype Repulsion Constraint to mitigate reward hacking in open-ended generation, and finally adapts the model to the Eastern cultural context with a small amount of supervision. Experimental results show that our method achieves stronger overall performance under the proposed evaluation framework, with particularly large gains in contextual fit and a better balance between image relevance and humor under cultural constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18091
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
Xu, Run
Li, Lu
Zhang, Rongzhao
Xu, Jie
Computation and Language
Computer Vision and Pattern Recognition
Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over explicit cultural context, making it difficult to jointly maintain image relevance, contextual appropriateness, and humor quality under a specified cultural background. To address this limitation, we introduce a new multimodal generation task, culture-aware humorous captioning, which requires a model to generate a humorous caption conditioned on both an input image and a target cultural context. Captions generated under different cultural contexts are not expected to share the same surface form, but should remain grounded in similar visual situations or humorous rationales.To support this task, we establish a six-dimensional evaluation framework covering image relevance, contextual fit, semantic richness, reasonableness, humor, and creativity. We further propose a staged alignment framework that first initializes the model with high-resource supervision under the Western cultural context, then performs multi-dimensional preference alignment via judge-based GRPO with a Degradation-aware Prototype Repulsion Constraint to mitigate reward hacking in open-ended generation, and finally adapts the model to the Eastern cultural context with a small amount of supervision. Experimental results show that our method achieves stronger overall performance under the proposed evaluation framework, with particularly large gains in contextual fit and a better balance between image relevance and humor under cultural constraints.
title Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.18091