State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pan, Zhaoyan, Li, Xiangdong, Wu, Wenke, Ma, Mengting, Lou, Ye, Zhou, Ji, Pan, Jiatong, Zhang, Wei
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913169918656512
author Pan, Zhaoyan
Li, Xiangdong
Wu, Wenke
Ma, Mengting
Lou, Ye
Zhou, Ji
Pan, Jiatong
Zhang, Wei
author_facet Pan, Zhaoyan
Li, Xiangdong
Wu, Wenke
Ma, Mengting
Lou, Ye
Zhou, Ji
Pan, Jiatong
Zhang, Wei
contents Conversational multimodal emotion recognition (MER) requires reliable prediction when language, acoustic, or visual observations are missing or unreliable. Many missing-modality methods reconstruct absent inputs, yet such recovery can be non-unique in dialogue context, and nonverbal cues may conflict with the target utterance. To this end, we propose CoRe-KD (Complete-view Reference-guided Knowledge Distillation), a state-anchored, conflict-regularized complete-view distillation framework for robust conversational MER. A complete-view teacher provides structured references, including prediction-level references, fused states, and modality-specific states. Complete-view State Anchoring (CSA) aligns incomplete-view student predictions and states with these references, while Nonverbal Conflict Exposure (NCE) trains on target-preserving nonverbal conflict views to reduce donor-label bias. Experiments on IEMOCAP and MELD, with CMU-MOSEI as a supplementary utterance-level check, show consistent gains under fixed- and random-missing protocols. Comprehensive ablation studies and further analyses support the role of CSA and the complementary effect of NCE.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29590
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
Pan, Zhaoyan
Li, Xiangdong
Wu, Wenke
Ma, Mengting
Lou, Ye
Zhou, Ji
Pan, Jiatong
Zhang, Wei
Multimedia
Conversational multimodal emotion recognition (MER) requires reliable prediction when language, acoustic, or visual observations are missing or unreliable. Many missing-modality methods reconstruct absent inputs, yet such recovery can be non-unique in dialogue context, and nonverbal cues may conflict with the target utterance. To this end, we propose CoRe-KD (Complete-view Reference-guided Knowledge Distillation), a state-anchored, conflict-regularized complete-view distillation framework for robust conversational MER. A complete-view teacher provides structured references, including prediction-level references, fused states, and modality-specific states. Complete-view State Anchoring (CSA) aligns incomplete-view student predictions and states with these references, while Nonverbal Conflict Exposure (NCE) trains on target-preserving nonverbal conflict views to reduce donor-label bias. Experiments on IEMOCAP and MELD, with CMU-MOSEI as a supplementary utterance-level check, show consistent gains under fixed- and random-missing protocols. Comprehensive ablation studies and further analyses support the role of CSA and the complementary effect of NCE.
title State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
topic Multimedia
url https://arxiv.org/abs/2605.29590