A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Praveen, R. Gnana, de Melo, Wheidima Carneiro, Ullah, Nasib, Aslam, Haseeb, Zeeshan, Osama, Denorme, Théo, Pedersoli, Marco, Koerich, Alessandro, Bacon, Simon, Cardinal, Patrick, Granger, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2021)
by: Praveen, R. Gnana, et al.
Published: (2021)
Joint Multimodal Transformer for Emotion Recognition in the Wild
by: Waligora, Paul, et al.
Published: (2024)
by: Waligora, Paul, et al.
Published: (2024)
Distilling Privileged Multimodal Information for Expression Recognition using Optimal Transport
by: Aslam, Muhammad Haseeb, et al.
Published: (2024)
by: Aslam, Muhammad Haseeb, et al.
Published: (2024)
Subject-Based Domain Adaptation for Facial Expression Recognition
by: Zeeshan, Muhammad Osama, et al.
Published: (2023)
by: Zeeshan, Muhammad Osama, et al.
Published: (2023)
Multi Teacher Privileged Knowledge Distillation for Multimodal Expression Recognition
by: Aslam, Muhammad Haseeb, et al.
Published: (2024)
by: Aslam, Muhammad Haseeb, et al.
Published: (2024)
Progressive Multi-Source Domain Adaptation for Personalized Facial Expression Recognition
by: Zeeshan, Muhammad Osama, et al.
Published: (2025)
by: Zeeshan, Muhammad Osama, et al.
Published: (2025)
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change
by: González-González, Manuela, et al.
Published: (2025)
by: González-González, Manuela, et al.
Published: (2025)
CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition
by: Zeeshan, Muhammad Osama, et al.
Published: (2026)
by: Zeeshan, Muhammad Osama, et al.
Published: (2026)
Deep Domain Adaptation for Ordinal Regression of Pain Intensity Estimation Using Weakly-Labelled Videos
by: Praveen, R. Gnana, et al.
Published: (2020)
by: Praveen, R. Gnana, et al.
Published: (2020)
Weakly Supervised Learning for Facial Affective Behavior Analysis : A Review
by: Praveen, R. Gnana, et al.
Published: (2021)
by: Praveen, R. Gnana, et al.
Published: (2021)
Deep Weakly-Supervised Domain Adaptation for Pain Localization in Videos
by: Praveen, R. Gnana, et al.
Published: (2019)
by: Praveen, R. Gnana, et al.
Published: (2019)
Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
by: González-González, Manuela, et al.
Published: (2026)
by: González-González, Manuela, et al.
Published: (2026)
Cross-Attention is Not Always Needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Disentangled Source-Free Personalization for Facial Expression Recognition with Neutral Target Data
by: Sharafi, Masoumeh, et al.
Published: (2025)
by: Sharafi, Masoumeh, et al.
Published: (2025)
Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild
by: Richet, Nicolas, et al.
Published: (2024)
by: Richet, Nicolas, et al.
Published: (2024)
Test-Time Adaptation via Cache Personalization for Facial Expression Recognition in Videos
by: Sharafi, Masoumeh, et al.
Published: (2026)
by: Sharafi, Masoumeh, et al.
Published: (2026)
MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
by: Zeeshan, Muhammad Osama, et al.
Published: (2025)
by: Zeeshan, Muhammad Osama, et al.
Published: (2025)
Guided Interpretable Facial Expression Recognition via Spatial Action Unit Cues
by: Belharbi, Soufiane, et al.
Published: (2024)
by: Belharbi, Soufiane, et al.
Published: (2024)
Spatial Action Unit Cues for Interpretable Deep Facial Expression Recognition
by: Belharbi, Soufiane, et al.
Published: (2024)
by: Belharbi, Soufiane, et al.
Published: (2024)
Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Personalized Feature Translation for Expression Recognition: An Efficient Source-Free Domain Adaptation Method
by: Sharafi, Masoumeh, et al.
Published: (2025)
by: Sharafi, Masoumeh, et al.
Published: (2025)
United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space
by: Praveen, R. Gnana, et al.
Published: (2025)
by: Praveen, R. Gnana, et al.
Published: (2025)
Learning from Stochastic Teacher Representations Using Student-Guided Knowledge Distillation
by: Aslam, Muhammad Haseeb, et al.
Published: (2025)
by: Aslam, Muhammad Haseeb, et al.
Published: (2025)
Dynamic Cross Attention for Audio-Visual Person Verification
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
by: Rajasekhar, Gnana Praveen, et al.
Published: (2025)
by: Rajasekhar, Gnana Praveen, et al.
Published: (2025)
Human Emotions Analysis and Recognition Using EEG Signals in Response to 360$^\circ$ Videos
by: Abbasi, Haseeb ur Rahman, et al.
Published: (2024)
by: Abbasi, Haseeb ur Rahman, et al.
Published: (2024)
Inconsistency-Aware Cross-Attention for Audio-Visual Fusion in Dimensional Emotion Recognition
by: Rajasekhar, G, et al.
Published: (2024)
by: Rajasekhar, G, et al.
Published: (2024)
From Affect to Complex Behavior: Advancing Multimodal Human-Centered AI at the 10th ABAW Workshop & Competition
by: Kollias, Dimitrios, et al.
Published: (2026)
by: Kollias, Dimitrios, et al.
Published: (2026)
A Realistic Protocol for Evaluation of Weakly Supervised Object Localization
by: Murtaza, Shakeeb, et al.
Published: (2024)
by: Murtaza, Shakeeb, et al.
Published: (2024)
Leveraging Transformers for Weakly Supervised Object Localization in Unconstrained Videos
by: Murtaza, Shakeeb, et al.
Published: (2024)
by: Murtaza, Shakeeb, et al.
Published: (2024)
ELMO: Efficiency via Low-precision and Peak Memory Optimization in Large Output Spaces
by: Zhang, Jinbin, et al.
Published: (2025)
by: Zhang, Jinbin, et al.
Published: (2025)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
by: Zhang, Jinbin, et al.
Published: (2025)
by: Zhang, Jinbin, et al.
Published: (2025)
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
by: Belal, Atif, et al.
Published: (2025)
by: Belal, Atif, et al.
Published: (2025)
MiPa: Mixed Patch Infrared-Visible Modality Agnostic Object Detection
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
Automatic Proficiency Assessment in L2 English Learners
by: Mohammadi, Armita, et al.
Published: (2025)
by: Mohammadi, Armita, et al.
Published: (2025)
Source-Free Domain Adaptation for YOLO Object Detection
by: Varailhon, Simon, et al.
Published: (2024)
by: Varailhon, Simon, et al.
Published: (2024)
Revisiting Mixout: An Overlooked Path to Robust Finetuning
by: Aminbeidokhti, Masih, et al.
Published: (2025)
by: Aminbeidokhti, Masih, et al.
Published: (2025)
A Review of Blockchain-based Smart Grid: Applications,Opportunities, and Future Directions
by: Ullah, H. Sami, et al.
Published: (2020)
by: Ullah, H. Sami, et al.
Published: (2020)
Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention
by: R, Joe Dhanith P, et al.
Published: (2024)
by: R, Joe Dhanith P, et al.
Published: (2024)
Similar Items
-
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2021) -
Joint Multimodal Transformer for Emotion Recognition in the Wild
by: Waligora, Paul, et al.
Published: (2024) -
Distilling Privileged Multimodal Information for Expression Recognition using Optimal Transport
by: Aslam, Muhammad Haseeb, et al.
Published: (2024) -
Subject-Based Domain Adaptation for Facial Expression Recognition
by: Zeeshan, Muhammad Osama, et al.
Published: (2023) -
Multi Teacher Privileged Knowledge Distillation for Multimodal Expression Recognition
by: Aslam, Muhammad Haseeb, et al.
Published: (2024)