Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Praveen, R. Gnana, Granger, Eric, Cardinal, Patrick |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2022)
by: Praveen, R. Gnana, et al.
Published: (2022)
Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space
by: Praveen, R. Gnana, et al.
Published: (2025)
by: Praveen, R. Gnana, et al.
Published: (2025)
Dynamic Cross Attention for Audio-Visual Person Verification
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
by: Chen, Xuanjun, et al.
Published: (2025)
by: Chen, Xuanjun, et al.
Published: (2025)
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
by: Rajasekhar, Gnana Praveen, et al.
Published: (2025)
by: Rajasekhar, Gnana Praveen, et al.
Published: (2025)
PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
by: Wen, Liuyuan
Published: (2024)
by: Wen, Liuyuan
Published: (2024)
Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes
by: Hein, Jonas, et al.
Published: (2025)
by: Hein, Jonas, et al.
Published: (2025)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
Machine Perceptual Quality: Evaluating the Impact of Severe Lossy Compression on Audio and Image Models
by: Jacobellis, Dan, et al.
Published: (2024)
by: Jacobellis, Dan, et al.
Published: (2024)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024)
by: Cao, Yuqin, et al.
Published: (2024)
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
by: Yeo, Jeong Hun, et al.
Published: (2023)
by: Yeo, Jeong Hun, et al.
Published: (2023)
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
by: Shimada, Kazuki, et al.
Published: (2025)
by: Shimada, Kazuki, et al.
Published: (2025)
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
by: Mingote, Victoria, et al.
Published: (2024)
by: Mingote, Victoria, et al.
Published: (2024)
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
by: Liu, Xiaofeng, et al.
Published: (2024)
by: Liu, Xiaofeng, et al.
Published: (2024)
Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Listening without Looking: Modality Bias in Audio-Visual Captioning
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
Efficient Face Detection with Audio-Based Region Proposals for Human-Robot Interactions
by: Aris, William, et al.
Published: (2023)
by: Aris, William, et al.
Published: (2023)
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
by: Yue, Xianghu, et al.
Published: (2024)
by: Yue, Xianghu, et al.
Published: (2024)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025)
by: Park, Jay, et al.
Published: (2025)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
by: Shou, Yuntao, et al.
Published: (2023)
by: Shou, Yuntao, et al.
Published: (2023)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
Spectral Bottleneck in Sinusoidal Representation Networks: Noise is All You Need
by: Chandravamsi, Hemanth, et al.
Published: (2025)
by: Chandravamsi, Hemanth, et al.
Published: (2025)
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
by: Jang, Youngjoon, et al.
Published: (2024)
by: Jang, Youngjoon, et al.
Published: (2024)
Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
by: Park, Se Jin, et al.
Published: (2023)
by: Park, Se Jin, et al.
Published: (2023)
Detection of Auditory Brainstem Response Peaks Using Image Processing Techniques in Infants with Normal Hearing Sensitivity
by: Majidpour, Amir, et al.
Published: (2024)
by: Majidpour, Amir, et al.
Published: (2024)
Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
by: Anand, et al.
Published: (2025)
by: Anand, et al.
Published: (2025)
Event2Audio: Event-Based Optical Vibration Sensing
by: Cai, Mingxuan, et al.
Published: (2025)
by: Cai, Mingxuan, et al.
Published: (2025)
PTSD-MDNN : Fusion tardive de réseaux de neurones profonds multimodaux pour la détection du trouble de stress post-traumatique
by: Nguyen-Phuoc, Long, et al.
Published: (2024)
by: Nguyen-Phuoc, Long, et al.
Published: (2024)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
by: Zhou, Huali, et al.
Published: (2024)
by: Zhou, Huali, et al.
Published: (2024)
Improvement Of Audiovisual Quality Estimation Using A Nonlinear Autoregressive Exogenous Neural Network And Bitstream Parameters
by: Kossi, Koffi, et al.
Published: (2024)
by: Kossi, Koffi, et al.
Published: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Ola: Pushing the Frontiers of Omni-Modal Language Model
by: Liu, Zuyan, et al.
Published: (2025)
by: Liu, Zuyan, et al.
Published: (2025)
Similar Items
-
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024) -
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2022) -
Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
by: Praveen, R. Gnana, et al.
Published: (2024) -
United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space
by: Praveen, R. Gnana, et al.
Published: (2025) -
Dynamic Cross Attention for Audio-Visual Person Verification
by: Praveen, R. Gnana, et al.
Published: (2024)