State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913169918656512 |
|---|---|
| author | Pan, Zhaoyan Li, Xiangdong Wu, Wenke Ma, Mengting Lou, Ye Zhou, Ji Pan, Jiatong Zhang, Wei |
| author_facet | Pan, Zhaoyan Li, Xiangdong Wu, Wenke Ma, Mengting Lou, Ye Zhou, Ji Pan, Jiatong Zhang, Wei |
| contents | Conversational multimodal emotion recognition (MER) requires reliable prediction when language, acoustic, or visual observations are missing or unreliable. Many missing-modality methods reconstruct absent inputs, yet such recovery can be non-unique in dialogue context, and nonverbal cues may conflict with the target utterance. To this end, we propose CoRe-KD (Complete-view Reference-guided Knowledge Distillation), a state-anchored, conflict-regularized complete-view distillation framework for robust conversational MER. A complete-view teacher provides structured references, including prediction-level references, fused states, and modality-specific states. Complete-view State Anchoring (CSA) aligns incomplete-view student predictions and states with these references, while Nonverbal Conflict Exposure (NCE) trains on target-preserving nonverbal conflict views to reduce donor-label bias. Experiments on IEMOCAP and MELD, with CMU-MOSEI as a supplementary utterance-level check, show consistent gains under fixed- and random-missing protocols. Comprehensive ablation studies and further analyses support the role of CSA and the complementary effect of NCE. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_29590 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition Pan, Zhaoyan Li, Xiangdong Wu, Wenke Ma, Mengting Lou, Ye Zhou, Ji Pan, Jiatong Zhang, Wei Multimedia Conversational multimodal emotion recognition (MER) requires reliable prediction when language, acoustic, or visual observations are missing or unreliable. Many missing-modality methods reconstruct absent inputs, yet such recovery can be non-unique in dialogue context, and nonverbal cues may conflict with the target utterance. To this end, we propose CoRe-KD (Complete-view Reference-guided Knowledge Distillation), a state-anchored, conflict-regularized complete-view distillation framework for robust conversational MER. A complete-view teacher provides structured references, including prediction-level references, fused states, and modality-specific states. Complete-view State Anchoring (CSA) aligns incomplete-view student predictions and states with these references, while Nonverbal Conflict Exposure (NCE) trains on target-preserving nonverbal conflict views to reduce donor-label bias. Experiments on IEMOCAP and MELD, with CMU-MOSEI as a supplementary utterance-level check, show consistent gains under fixed- and random-missing protocols. Comprehensive ablation studies and further analyses support the role of CSA and the complementary effect of NCE. |
| title | State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition |
| topic | Multimedia |
| url | https://arxiv.org/abs/2605.29590 |