Multi-View Incongruity Learning for Multimodal Sarcasm Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Diandian, Cao, Cong, Yuan, Fangfang, Liu, Yanbing, Zeng, Guangjie, Yu, Xiaoyan, Peng, Hao, Yu, Philip S.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912147644088320
author Guo, Diandian
Cao, Cong
Yuan, Fangfang
Liu, Yanbing
Zeng, Guangjie
Yu, Xiaoyan
Peng, Hao
Yu, Philip S.
author_facet Guo, Diandian
Cao, Cong
Yuan, Fangfang
Liu, Yanbing
Zeng, Guangjie
Yu, Xiaoyan
Peng, Hao
Yu, Philip S.
contents Multimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions, demonstrating poor generalizability beyond training environments. Regarding this phenomenon, this paper undertakes several initiatives. Firstly, we identify two primary causes that lead to the reliance of spurious correlations. Secondly, we address these challenges by proposing a novel method that integrate Multimodal Incongruities via Contrastive Learning (MICL) for multimodal sarcasm detection. Specifically, we first leverage incongruity to drive multi-view learning from three views: token-patch, entity-object, and sentiment. Then, we introduce extensive data augmentation to mitigate the biased learning of the textual modality. Additionally, we construct a test set, SPMSD, which consists potential spurious correlations to evaluate the the model's generalizability. Experimental results demonstrate the superiority of MICL on benchmark datasets, along with the analyses showcasing MICL's advancement in mitigating the effect of spurious correlation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00756
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-View Incongruity Learning for Multimodal Sarcasm Detection
Guo, Diandian
Cao, Cong
Yuan, Fangfang
Liu, Yanbing
Zeng, Guangjie
Yu, Xiaoyan
Peng, Hao
Yu, Philip S.
Computation and Language
Multimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions, demonstrating poor generalizability beyond training environments. Regarding this phenomenon, this paper undertakes several initiatives. Firstly, we identify two primary causes that lead to the reliance of spurious correlations. Secondly, we address these challenges by proposing a novel method that integrate Multimodal Incongruities via Contrastive Learning (MICL) for multimodal sarcasm detection. Specifically, we first leverage incongruity to drive multi-view learning from three views: token-patch, entity-object, and sentiment. Then, we introduce extensive data augmentation to mitigate the biased learning of the textual modality. Additionally, we construct a test set, SPMSD, which consists potential spurious correlations to evaluate the the model's generalizability. Experimental results demonstrate the superiority of MICL on benchmark datasets, along with the analyses showcasing MICL's advancement in mitigating the effect of spurious correlation.
title Multi-View Incongruity Learning for Multimodal Sarcasm Detection
topic Computation and Language
url https://arxiv.org/abs/2412.00756