Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Zhaoyan, Zhou, Hengyang, Li, Xiangdong, Wang, Yuning, Lou, Ye, Pan, Jiatong, Zhou, Ji, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
von: Wang, Jiadong, et al.
Veröffentlicht: (2026)
von: Wang, Jiadong, et al.
Veröffentlicht: (2026)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Towards Multimodal Emotional Support Conversation Systems
von: Chu, Yuqi, et al.
Veröffentlicht: (2024)
von: Chu, Yuqi, et al.
Veröffentlicht: (2024)
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
AcoustEmo: Open-Vocabulary Emotion Reasoning via Utterance-Aware Acoustic Q-Former
von: Zhang, Liyun, et al.
Veröffentlicht: (2026)
von: Zhang, Liyun, et al.
Veröffentlicht: (2026)
Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis
von: Dong, Guangyuan, et al.
Veröffentlicht: (2026)
von: Dong, Guangyuan, et al.
Veröffentlicht: (2026)
MInD: Improving Multimodal Sentiment Analysis via Multimodal Information Disentanglement
von: Dai, Weichen, et al.
Veröffentlicht: (2024)
von: Dai, Weichen, et al.
Veröffentlicht: (2024)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Multi-source Knowledge Enhanced Graph Attention Networks for Multimodal Fact Verification
von: Cao, Han, et al.
Veröffentlicht: (2024)
von: Cao, Han, et al.
Veröffentlicht: (2024)
PRISM: Exposing and Resolving Spurious Isolation in Federated Multimodal Continual Learning
von: Wu, Beining, et al.
Veröffentlicht: (2026)
von: Wu, Beining, et al.
Veröffentlicht: (2026)
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning
von: Xiong, Zechang, et al.
Veröffentlicht: (2026)
von: Xiong, Zechang, et al.
Veröffentlicht: (2026)
Multimodal Graph-Based Variational Mixture of Experts Network for Zero-Shot Multimodal Information Extraction
von: Zhou, Baohang, et al.
Veröffentlicht: (2025)
von: Zhou, Baohang, et al.
Veröffentlicht: (2025)
Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
von: Huang, Yiheng, et al.
Veröffentlicht: (2025)
von: Huang, Yiheng, et al.
Veröffentlicht: (2025)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Scaling up Multimodal Pre-training for Sign Language Understanding
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition with Large Language Models
von: Zhang, Hongrui, et al.
Veröffentlicht: (2026)
von: Zhang, Hongrui, et al.
Veröffentlicht: (2026)
Angle-Optimized Partial Disentanglement for Multimodal Emotion Recognition in Conversation
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
von: An, Xiao, et al.
Veröffentlicht: (2026)
von: An, Xiao, et al.
Veröffentlicht: (2026)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2026)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2026)
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2023)
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2023)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
von: Feng, X., et al.
Veröffentlicht: (2024)
von: Feng, X., et al.
Veröffentlicht: (2024)
Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
Beyond Static Collision Handling: Adaptive Semantic ID Learning for Multimodal Recommendation at Industrial Scale
von: Pan, Yongsen, et al.
Veröffentlicht: (2026)
von: Pan, Yongsen, et al.
Veröffentlicht: (2026)
Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion Recognition
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2024)
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026) -
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025) -
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024) -
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
von: Wang, Jiadong, et al.
Veröffentlicht: (2026) -
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)