Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Praveen, R. Gnana, Alam, Jahangir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
Dynamic Cross Attention for Audio-Visual Person Verification
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2025)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2025)
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2021)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2021)
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2022)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2022)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
von: Shou, Yuntao, et al.
Veröffentlicht: (2023)
von: Shou, Yuntao, et al.
Veröffentlicht: (2023)
Joint Multimodal Transformer for Emotion Recognition in the Wild
von: Waligora, Paul, et al.
Veröffentlicht: (2024)
von: Waligora, Paul, et al.
Veröffentlicht: (2024)
Dynamic Modality and View Selection for Multimodal Emotion Recognition with Missing Modalities
von: Menon, Luciana Trinkaus, et al.
Veröffentlicht: (2024)
von: Menon, Luciana Trinkaus, et al.
Veröffentlicht: (2024)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
von: Liu, Lei, et al.
Veröffentlicht: (2024)
von: Liu, Lei, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
von: Jiang, Peiyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Peiyuan, et al.
Veröffentlicht: (2025)
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
von: Zhao, Fuzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Fuzheng, et al.
Veröffentlicht: (2024)
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
von: D., Quang-Anh N., et al.
Veröffentlicht: (2024)
von: D., Quang-Anh N., et al.
Veröffentlicht: (2024)
Decoding Emotions: Unveiling Facial Expressions through Acoustic Sensing with Contrastive Attention
von: Wang, Guangjing, et al.
Veröffentlicht: (2024)
von: Wang, Guangjing, et al.
Veröffentlicht: (2024)
Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
von: Ortega-Beltrán, Elena, et al.
Veröffentlicht: (2024)
von: Ortega-Beltrán, Elena, et al.
Veröffentlicht: (2024)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025)
von: Anand, et al.
Veröffentlicht: (2025)
Multimodal Fusion Method with Spatiotemporal Sequences and Relationship Learning for Valence-Arousal Estimation
von: Yu, Jun, et al.
Veröffentlicht: (2024)
von: Yu, Jun, et al.
Veröffentlicht: (2024)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Emotional Face-to-Speech
von: Ye, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ye, Jiaxin, et al.
Veröffentlicht: (2025)
Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
von: Guragain, Anmol
Veröffentlicht: (2026)
von: Guragain, Anmol
Veröffentlicht: (2026)
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Cross-Attention is Not Always Needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024)
Digit Recognition using Multimodal Spiking Neural Networks
von: Bjorndahl, William, et al.
Veröffentlicht: (2024)
von: Bjorndahl, William, et al.
Veröffentlicht: (2024)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2024)
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2024)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
von: Katamneni, Vinaya Sree, et al.
Veröffentlicht: (2024)
von: Katamneni, Vinaya Sree, et al.
Veröffentlicht: (2024)
AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations
von: Devulapally, Naresh Kumar, et al.
Veröffentlicht: (2024)
von: Devulapally, Naresh Kumar, et al.
Veröffentlicht: (2024)
Exploring Multi-Modal Control in Music-Driven Dance Generation
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
Dual Audio-Centric Modality Coupling for Talking Head Generation
von: Fu, Ao, et al.
Veröffentlicht: (2025)
von: Fu, Ao, et al.
Veröffentlicht: (2025)
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
von: Kwon, Mingi, et al.
Veröffentlicht: (2025)
von: Kwon, Mingi, et al.
Veröffentlicht: (2025)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
von: Fei, Zhengcong, et al.
Veröffentlicht: (2023)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024) -
Dynamic Cross Attention for Audio-Visual Person Verification
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2024) -
United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2025) -
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2021) -
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)