URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhenyu, Cheng, Weichen, Li, Weijia, Mou, Junjie, Zhao, Zongyou, Zhang, Guoying
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910187005149184
author Wang, Zhenyu
Cheng, Weichen
Li, Weijia
Mou, Junjie
Zhao, Zongyou
Zhang, Guoying
author_facet Wang, Zhenyu
Cheng, Weichen
Li, Weijia
Mou, Junjie
Zhao, Zongyou
Zhang, Guoying
contents Multimodal sarcasm detection (MSD) aims to identify sarcastic intent from semantic incongruity between text and image. Although recent methods have improved MSD through cross-modal interaction and incongruity reasoning, most still treat modalities as equally reliable. In real social media posts, however, text and images often differ in noise level and relevance, making deterministic fusion susceptible to noisy evidence and weakened incongruity cues. To address this issue, we propose Uncertainty-aware Robust Multimodal Fusion (URMF), a unified framework for robust MSD. URMF first injects visual evidence into textual representations through multi-head cross-attention, and then applies self-attention in the fused semantic space to enhance incongruity reasoning. It models textual, visual, and interaction-aware representations as learnable Gaussian posteriors to estimate modality-specific uncertainty. Based on the estimated uncertainty, URMF dynamically adjusts modality contributions during fusion to suppress unreliable evidence. We further optimize the model with a unified objective that combines information bottleneck regularization, modality prior regularization, cross-modal distribution alignment, and uncertainty-driven contrastive learning. Experiments on the public MSD and MMSD2 benchmarks show that URMF outperforms representative unimodal, multimodal, and MLLM-based baselines. The results demonstrate that explicit uncertainty modeling can improve both accuracy and robustness in multimodal sarcasm detection.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06728
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
Wang, Zhenyu
Cheng, Weichen
Li, Weijia
Mou, Junjie
Zhao, Zongyou
Zhang, Guoying
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Multimodal sarcasm detection (MSD) aims to identify sarcastic intent from semantic incongruity between text and image. Although recent methods have improved MSD through cross-modal interaction and incongruity reasoning, most still treat modalities as equally reliable. In real social media posts, however, text and images often differ in noise level and relevance, making deterministic fusion susceptible to noisy evidence and weakened incongruity cues. To address this issue, we propose Uncertainty-aware Robust Multimodal Fusion (URMF), a unified framework for robust MSD. URMF first injects visual evidence into textual representations through multi-head cross-attention, and then applies self-attention in the fused semantic space to enhance incongruity reasoning. It models textual, visual, and interaction-aware representations as learnable Gaussian posteriors to estimate modality-specific uncertainty. Based on the estimated uncertainty, URMF dynamically adjusts modality contributions during fusion to suppress unreliable evidence. We further optimize the model with a unified objective that combines information bottleneck regularization, modality prior regularization, cross-modal distribution alignment, and uncertainty-driven contrastive learning. Experiments on the public MSD and MMSD2 benchmarks show that URMF outperforms representative unimodal, multimodal, and MLLM-based baselines. The results demonstrate that explicit uncertainty modeling can improve both accuracy and robustness in multimodal sarcasm detection.
title URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2604.06728