Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhu, Gao, Xiyuan, Zhang, Yuqing, Nayak, Shekhar, Coler, Matt
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908546592931840
author Li, Zhu
Gao, Xiyuan
Zhang, Yuqing
Nayak, Shekhar
Coler, Matt
author_facet Li, Zhu
Gao, Xiyuan
Zhang, Yuqing
Nayak, Shekhar
Coler, Matt
contents Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior work has primarily focused on textual or visual-textual sarcasm, comprehensive audio-visual-textual sarcasm understanding remains underexplored. In this paper, we systematically evaluate large language models (LLMs) and multimodal LLMs for sarcasm detection on English (MUStARD++) and Chinese (MCSD 1.0) in zero-shot, few-shot, and LoRA fine-tuning settings. In addition to direct classification, we explore models as feature encoders, integrating their representations through a collaborative gating fusion module. Experimental results show that audio-based models achieve the strongest unimodal performance, while text-audio and audio-vision combinations outperform unimodal and trimodal models. Furthermore, MLLMs such as Qwen-Omni show competitive zero-shot and fine-tuned performance. Our findings highlight the potential of MLLMs for cross-lingual, audio-visual-textual sarcasm understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
Li, Zhu
Gao, Xiyuan
Zhang, Yuqing
Nayak, Shekhar
Coler, Matt
Computation and Language
Multimedia
Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior work has primarily focused on textual or visual-textual sarcasm, comprehensive audio-visual-textual sarcasm understanding remains underexplored. In this paper, we systematically evaluate large language models (LLMs) and multimodal LLMs for sarcasm detection on English (MUStARD++) and Chinese (MCSD 1.0) in zero-shot, few-shot, and LoRA fine-tuning settings. In addition to direct classification, we explore models as feature encoders, integrating their representations through a collaborative gating fusion module. Experimental results show that audio-based models achieve the strongest unimodal performance, while text-audio and audio-vision combinations outperform unimodal and trimodal models. Furthermore, MLLMs such as Qwen-Omni show competitive zero-shot and fine-tuned performance. Our findings highlight the potential of MLLMs for cross-lingual, audio-visual-textual sarcasm understanding.
title Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2509.15476