Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Qianrui, Xu, Hua, Li, Hao, Zhang, Hanlei, Zhang, Xiaohan, Wang, Yifan, Gao, Kai
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910473641787392
author Zhou, Qianrui
Xu, Hua
Li, Hao
Zhang, Hanlei
Zhang, Xiaohan
Wang, Yifan
Gao, Kai
author_facet Zhou, Qianrui
Xu, Hua
Li, Hao
Zhang, Hanlei
Zhang, Xiaohan
Wang, Yifan
Gao, Kai
contents Multimodal intent recognition aims to leverage diverse modalities such as expressions, body movements and tone of speech to comprehend user's intent, constituting a critical task for understanding human language and behavior in real-world multimodal scenarios. Nevertheless, the majority of existing methods ignore potential correlations among different modalities and own limitations in effectively learning semantic features from nonverbal modalities. In this paper, we introduce a token-level contrastive learning method with modality-aware prompting (TCL-MAP) to address the above challenges. To establish an optimal multimodal semantic environment for text modality, we develop a modality-aware prompting module (MAP), which effectively aligns and fuses features from text, video and audio modalities with similarity-based modality alignment and cross-modality attention mechanism. Based on the modality-aware prompt and ground truth labels, the proposed token-level contrastive learning framework (TCL) constructs augmented samples and employs NT-Xent loss on the label token. Specifically, TCL capitalizes on the optimal textual semantic insights derived from intent labels to guide the learning processes of other modalities in return. Extensive experiments show that our method achieves remarkable improvements compared to state-of-the-art methods. Additionally, ablation analyses demonstrate the superiority of the modality-aware prompt over the handcrafted prompt, which holds substantial significance for multimodal prompt learning. The codes are released at https://github.com/thuiar/TCL-MAP.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14667
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
Zhou, Qianrui
Xu, Hua
Li, Hao
Zhang, Hanlei
Zhang, Xiaohan
Wang, Yifan
Gao, Kai
Multimedia
Machine Learning
Multimodal intent recognition aims to leverage diverse modalities such as expressions, body movements and tone of speech to comprehend user's intent, constituting a critical task for understanding human language and behavior in real-world multimodal scenarios. Nevertheless, the majority of existing methods ignore potential correlations among different modalities and own limitations in effectively learning semantic features from nonverbal modalities. In this paper, we introduce a token-level contrastive learning method with modality-aware prompting (TCL-MAP) to address the above challenges. To establish an optimal multimodal semantic environment for text modality, we develop a modality-aware prompting module (MAP), which effectively aligns and fuses features from text, video and audio modalities with similarity-based modality alignment and cross-modality attention mechanism. Based on the modality-aware prompt and ground truth labels, the proposed token-level contrastive learning framework (TCL) constructs augmented samples and employs NT-Xent loss on the label token. Specifically, TCL capitalizes on the optimal textual semantic insights derived from intent labels to guide the learning processes of other modalities in return. Extensive experiments show that our method achieves remarkable improvements compared to state-of-the-art methods. Additionally, ablation analyses demonstrate the superiority of the modality-aware prompt over the handcrafted prompt, which holds substantial significance for multimodal prompt learning. The codes are released at https://github.com/thuiar/TCL-MAP.
title Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
topic Multimedia
Machine Learning
url https://arxiv.org/abs/2312.14667