Deep Learning Approaches for Multimodal Intent Recognition: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Jingwei, Wen, Yuhua, Li, Qifei, Hu, Minchi, Zhou, Yingying, Xue, Jingyao, Wu, Junyang, Gao, Yingming, Wen, Zhengqi, Tao, Jianhua, Li, Ya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
by: Wen, Yuhua, et al.
Published: (2025)
by: Wen, Yuhua, et al.
Published: (2025)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
by: Li, Qifei, et al.
Published: (2024)
by: Li, Qifei, et al.
Published: (2024)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
Spatio-Temporal Cluster-Triggered Encoding for Spiking Neural Networks
by: Hu, Minchi
Published: (2025)
by: Hu, Minchi
Published: (2025)
Psy-Insight: Explainable Multi-turn Bilingual Dataset for Mental Health Counseling
by: Chen, Keqi, et al.
Published: (2025)
by: Chen, Keqi, et al.
Published: (2025)
Exploring the Role of Audio in Multimodal Misinformation Detection
by: Liu, Moyang, et al.
Published: (2024)
by: Liu, Moyang, et al.
Published: (2024)
PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis
by: Xie, Heng, et al.
Published: (2025)
by: Xie, Heng, et al.
Published: (2025)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
by: Ren, Yong, et al.
Published: (2026)
by: Ren, Yong, et al.
Published: (2026)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
by: Xue, Jinlong, et al.
Published: (2024)
by: Xue, Jinlong, et al.
Published: (2024)
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
by: Xue, Jinlong, et al.
Published: (2024)
by: Xue, Jinlong, et al.
Published: (2024)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
COAL: Robust Contrastive Learning‐Based Visual Navigation Framework
by: Zengmao Wang, et al.
Published: (2025)
by: Zengmao Wang, et al.
Published: (2025)
Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning
by: Yan, Kaiying, et al.
Published: (2025)
by: Yan, Kaiying, et al.
Published: (2025)
Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement
by: Yang, Zhengxian, et al.
Published: (2026)
by: Yang, Zhengxian, et al.
Published: (2026)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2023)
by: Zhou, Qianrui, et al.
Published: (2023)
A Survey on Backbones for Deep Video Action Recognition
by: Tang, Zixuan, et al.
Published: (2024)
by: Tang, Zixuan, et al.
Published: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
by: Xue, Jinlong, et al.
Published: (2024)
by: Xue, Jinlong, et al.
Published: (2024)
Deep Smart Contract Intent Detection
by: Huang, Youwei, et al.
Published: (2022)
by: Huang, Youwei, et al.
Published: (2022)
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
by: Zhang, Haojie, et al.
Published: (2024)
by: Zhang, Haojie, et al.
Published: (2024)
Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
Fake News Detection and Manipulation Reasoning via Large Vision-Language Models
by: Jin, Ruihan, et al.
Published: (2024)
by: Jin, Ruihan, et al.
Published: (2024)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026)
by: Zhou, Qianrui, et al.
Published: (2026)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
by: Fan, Cunhang, et al.
Published: (2023)
by: Fan, Cunhang, et al.
Published: (2023)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
by: Bai, Bingsong, et al.
Published: (2024)
by: Bai, Bingsong, et al.
Published: (2024)
Residual Speaker Representation for One-Shot Voice Conversion
by: Xu, Le, et al.
Published: (2023)
by: Xu, Le, et al.
Published: (2023)
AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
by: Zhou, Junzuo, et al.
Published: (2024)
by: Zhou, Junzuo, et al.
Published: (2024)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2025)
by: Zhou, Qianrui, et al.
Published: (2025)
MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
by: Yan, Kaiying, et al.
Published: (2025)
by: Yan, Kaiying, et al.
Published: (2025)
TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression Recognition
by: Zhu, Jianhua, et al.
Published: (2024)
by: Zhu, Jianhua, et al.
Published: (2024)
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
by: Zhang, Zhen, et al.
Published: (2024)
by: Zhang, Zhen, et al.
Published: (2024)
A Heuristic Model for Improving Physical Education in the Higher Education System
by: Jingwei Wen
Published: (2025)
by: Jingwei Wen
Published: (2025)
WDMIR: Wavelet-Driven Multimodal Intent Recognition
by: Gong, Weiyin, et al.
Published: (2025)
by: Gong, Weiyin, et al.
Published: (2025)
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
by: Jin, Ruihan, et al.
Published: (2025)
by: Jin, Ruihan, et al.
Published: (2025)
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
by: Shou, Yuntao, et al.
Published: (2025)
by: Shou, Yuntao, et al.
Published: (2025)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
Similar Items
-
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
by: Wen, Yuhua, et al.
Published: (2025) -
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
by: Li, Qifei, et al.
Published: (2024) -
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
by: Wang, Cong, et al.
Published: (2025) -
Spatio-Temporal Cluster-Triggered Encoding for Spiking Neural Networks
by: Hu, Minchi
Published: (2025) -
Psy-Insight: Explainable Multi-turn Bilingual Dataset for Mental Health Counseling
by: Chen, Keqi, et al.
Published: (2025)