WDMIR: Wavelet-Driven Multimodal Intent Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Weiyin, Zhang, Kai, Zhang, Yanghai, Liu, Qi, Sun, Xinjie, Lu, Junyu, Zhu, Linbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908405600354304
author Gong, Weiyin
Zhang, Kai
Zhang, Yanghai
Liu, Qi
Sun, Xinjie
Lu, Junyu
Zhu, Linbo
author_facet Gong, Weiyin
Zhang, Kai
Zhang, Yanghai
Liu, Qi
Sun, Xinjie
Lu, Junyu
Zhu, Linbo
contents Multimodal intent recognition (MIR) seeks to accurately interpret user intentions by integrating verbal and non-verbal information across video, audio and text modalities. While existing approaches prioritize text analysis, they often overlook the rich semantic content embedded in non-verbal cues. This paper presents a novel Wavelet-Driven Multimodal Intent Recognition(WDMIR) framework that enhances intent understanding through frequency-domain analysis of non-verbal information. To be more specific, we propose: (1) a wavelet-driven fusion module that performs synchronized decomposition and integration of video-audio features in the frequency domain, enabling fine-grained analysis of temporal dynamics; (2) a cross-modal interaction mechanism that facilitates progressive feature enhancement from bimodal to trimodal integration, effectively bridging the semantic gap between verbal and non-verbal information. Extensive experiments on MIntRec demonstrate that our approach achieves state-of-the-art performance, surpassing previous methods by 1.13% on accuracy. Ablation studies further verify that the wavelet-driven fusion module significantly improves the extraction of semantic information from non-verbal sources, with a 0.41% increase in recognition accuracy when analyzing subtle emotional cues.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10011
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WDMIR: Wavelet-Driven Multimodal Intent Recognition
Gong, Weiyin
Zhang, Kai
Zhang, Yanghai
Liu, Qi
Sun, Xinjie
Lu, Junyu
Zhu, Linbo
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Signal Processing
Multimodal intent recognition (MIR) seeks to accurately interpret user intentions by integrating verbal and non-verbal information across video, audio and text modalities. While existing approaches prioritize text analysis, they often overlook the rich semantic content embedded in non-verbal cues. This paper presents a novel Wavelet-Driven Multimodal Intent Recognition(WDMIR) framework that enhances intent understanding through frequency-domain analysis of non-verbal information. To be more specific, we propose: (1) a wavelet-driven fusion module that performs synchronized decomposition and integration of video-audio features in the frequency domain, enabling fine-grained analysis of temporal dynamics; (2) a cross-modal interaction mechanism that facilitates progressive feature enhancement from bimodal to trimodal integration, effectively bridging the semantic gap between verbal and non-verbal information. Extensive experiments on MIntRec demonstrate that our approach achieves state-of-the-art performance, surpassing previous methods by 1.13% on accuracy. Ablation studies further verify that the wavelet-driven fusion module significantly improves the extraction of semantic information from non-verbal sources, with a 0.41% increase in recognition accuracy when analyzing subtle emotional cues.
title WDMIR: Wavelet-Driven Multimodal Intent Recognition
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Signal Processing
url https://arxiv.org/abs/2506.10011