Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Wen-Jue, Zhu, Xiaofeng, Zhang, Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918264362237952
author He, Wen-Jue
Zhu, Xiaofeng
Zhang, Zheng
author_facet He, Wen-Jue
Zhu, Xiaofeng
Zhang, Zheng
contents Incomplete multi-modal emotion recognition (IMER) aims at understanding human intentions and sentiments by comprehensively exploring the partially observed multi-source data. Although the multi-modal data is expected to provide more abundant information, the performance gap and modality under-optimization problem hinder effective multi-modal learning in practice, and are exacerbated in the confrontation of the missing data. To address this issue, we devise a novel Cross-modal Prompting (ComP) method, which emphasizes coherent information by enhancing modality-specific features and improves the overall recognition accuracy by boosting each modality's performance. Specifically, a progressive prompt generation module with a dynamic gradient modulator is proposed to produce concise and consistent modality semantic cues. Meanwhile, cross-modal knowledge propagation selectively amplifies the consistent information in modality features with the delivered prompts to enhance the discrimination of the modality-specific output. Additionally, a coordinator is designed to dynamically re-weight the modality outputs as a complement to the balance strategy to improve the model's efficacy. Extensive experiments on 4 datasets with 7 SOTA methods under different missing rates validate the effectiveness of our proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11239
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
He, Wen-Jue
Zhu, Xiaofeng
Zhang, Zheng
Computer Vision and Pattern Recognition
Incomplete multi-modal emotion recognition (IMER) aims at understanding human intentions and sentiments by comprehensively exploring the partially observed multi-source data. Although the multi-modal data is expected to provide more abundant information, the performance gap and modality under-optimization problem hinder effective multi-modal learning in practice, and are exacerbated in the confrontation of the missing data. To address this issue, we devise a novel Cross-modal Prompting (ComP) method, which emphasizes coherent information by enhancing modality-specific features and improves the overall recognition accuracy by boosting each modality's performance. Specifically, a progressive prompt generation module with a dynamic gradient modulator is proposed to produce concise and consistent modality semantic cues. Meanwhile, cross-modal knowledge propagation selectively amplifies the consistent information in modality features with the delivered prompts to enhance the discrimination of the modality-specific output. Additionally, a coordinator is designed to dynamically re-weight the modality outputs as a complement to the balance strategy to improve the model's efficacy. Extensive experiments on 4 datasets with 7 SOTA methods under different missing rates validate the effectiveness of our proposed method.
title Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11239