Can Large Vision-Language Models Understand Multimodal Sarcasm?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xinyu, Zhang, Yue, Jing, Liqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915428949819392
author Wang, Xinyu
Zhang, Yue
Jing, Liqiang
author_facet Wang, Xinyu
Zhang, Yue
Jing, Liqiang
contents Sarcasm is a complex linguistic phenomenon that involves a disparity between literal and intended meanings, making it challenging for sentiment analysis and other emotion-sensitive tasks. While traditional sarcasm detection methods primarily focus on text, recent approaches have incorporated multimodal information. However, the application of Large Visual Language Models (LVLMs) in Multimodal Sarcasm Analysis (MSA) remains underexplored. In this paper, we evaluate LVLMs in MSA tasks, specifically focusing on Multimodal Sarcasm Detection and Multimodal Sarcasm Explanation. Through comprehensive experiments, we identify key limitations, such as insufficient visual understanding and a lack of conceptual knowledge. To address these issues, we propose a training-free framework that integrates in-depth object extraction and external conceptual knowledge to improve the model's ability to interpret and explain sarcasm in multimodal contexts. The experimental results on multiple models show the effectiveness of our proposed framework. The code is available at https://github.com/cp-cp/LVLM-MSA.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03654
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Large Vision-Language Models Understand Multimodal Sarcasm?
Wang, Xinyu
Zhang, Yue
Jing, Liqiang
Computation and Language
Computer Vision and Pattern Recognition
Sarcasm is a complex linguistic phenomenon that involves a disparity between literal and intended meanings, making it challenging for sentiment analysis and other emotion-sensitive tasks. While traditional sarcasm detection methods primarily focus on text, recent approaches have incorporated multimodal information. However, the application of Large Visual Language Models (LVLMs) in Multimodal Sarcasm Analysis (MSA) remains underexplored. In this paper, we evaluate LVLMs in MSA tasks, specifically focusing on Multimodal Sarcasm Detection and Multimodal Sarcasm Explanation. Through comprehensive experiments, we identify key limitations, such as insufficient visual understanding and a lack of conceptual knowledge. To address these issues, we propose a training-free framework that integrates in-depth object extraction and external conceptual knowledge to improve the model's ability to interpret and explain sarcasm in multimodal contexts. The experimental results on multiple models show the effectiveness of our proposed framework. The code is available at https://github.com/cp-cp/LVLM-MSA.
title Can Large Vision-Language Models Understand Multimodal Sarcasm?
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.03654