Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Pendyala, Varsha, Morgado, Pedro, Sethares, William |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
by: Yeo, Jeong Hun, et al.
Published: (2023)
by: Yeo, Jeong Hun, et al.
Published: (2023)
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
by: Salaj, Ina, et al.
Published: (2025)
by: Salaj, Ina, et al.
Published: (2025)
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023)
by: Premananth, Gowtham, et al.
Published: (2023)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024)
by: Cao, Yuqin, et al.
Published: (2024)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
by: Mingote, Victoria, et al.
Published: (2024)
by: Mingote, Victoria, et al.
Published: (2024)
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
by: Yue, Xianghu, et al.
Published: (2024)
by: Yue, Xianghu, et al.
Published: (2024)
LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation
by: Jacobellis, Dan, et al.
Published: (2026)
by: Jacobellis, Dan, et al.
Published: (2026)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
by: Ragano, Alessandro, et al.
Published: (2024)
by: Ragano, Alessandro, et al.
Published: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
by: Chou, Huang-Cheng, et al.
Published: (2024)
by: Chou, Huang-Cheng, et al.
Published: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2024)
by: Berghi, Davide, et al.
Published: (2024)
On the Parameter Estimation of Sinusoidal Models for Speech and Audio Signals
by: Kafentzis, George P.
Published: (2024)
by: Kafentzis, George P.
Published: (2024)
Multimodal sensor fusion for real-time location-dependent defect detection in laser-directed energy deposition
by: Chen, Lequn, et al.
Published: (2023)
by: Chen, Lequn, et al.
Published: (2023)
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
by: Hopkins, Torin, et al.
Published: (2026)
by: Hopkins, Torin, et al.
Published: (2026)
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
by: Yeo, Jeong Hun, et al.
Published: (2025)
by: Yeo, Jeong Hun, et al.
Published: (2025)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024)
by: Cai, Zhuojiang, et al.
Published: (2024)
PersonaCite: VoC-Grounded Interviewable Agentic Synthetic AI Personas for Verifiable User and Design Research
by: Truss, Mario
Published: (2026)
by: Truss, Mario
Published: (2026)
Multimodal Machine Learning Can Predict Videoconference Fluidity and Enjoyment
by: Chang, Andrew, et al.
Published: (2025)
by: Chang, Andrew, et al.
Published: (2025)
Lessons Learned from Developing a Privacy-Preserving Multimodal Wearable for Local Voice-and-Vision Inference
by: Tussa, Yonatan, et al.
Published: (2025)
by: Tussa, Yonatan, et al.
Published: (2025)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
by: Sun, Licai, et al.
Published: (2024)
by: Sun, Licai, et al.
Published: (2024)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
by: Eliav, Amit, et al.
Published: (2024)
by: Eliav, Amit, et al.
Published: (2024)
Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
by: Wang, Zanxu, et al.
Published: (2025)
by: Wang, Zanxu, et al.
Published: (2025)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
by: Sulun, Serkan, et al.
Published: (2025)
by: Sulun, Serkan, et al.
Published: (2025)
Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
by: Wen, Liuyuan
Published: (2024)
by: Wen, Liuyuan
Published: (2024)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
by: Xu, Zhongweiyang, et al.
Published: (2024)
by: Xu, Zhongweiyang, et al.
Published: (2024)
Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios with Synthetic Visual Data
by: Buitrago, Pol, et al.
Published: (2026)
by: Buitrago, Pol, et al.
Published: (2026)
Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
by: Liu, Xiaofeng, et al.
Published: (2024)
by: Liu, Xiaofeng, et al.
Published: (2024)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
by: Yu, Fan, et al.
Published: (2024)
by: Yu, Fan, et al.
Published: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
by: Liu, Qianhui, et al.
Published: (2024)
by: Liu, Qianhui, et al.
Published: (2024)
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
by: Hnatyshyn, Rostyslav, et al.
Published: (2024)
by: Hnatyshyn, Rostyslav, et al.
Published: (2024)
Assessing the Viability of Wave Field Synthesis in VR-Based Cognitive Research
by: Kahl, Benjamin
Published: (2025)
by: Kahl, Benjamin
Published: (2025)
Creating Aesthetic Sonifications on the Web with SIREN
by: Peng, Tristan, et al.
Published: (2024)
by: Peng, Tristan, et al.
Published: (2024)
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
by: Huh, Mina, et al.
Published: (2026)
by: Huh, Mina, et al.
Published: (2026)
NeoLightning: A Modern Reimagination of Gesture-Based Sound Design
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
by: Zang, Yongyi, et al.
Published: (2023)
by: Zang, Yongyi, et al.
Published: (2023)
W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling
by: Babaei, Hossein, et al.
Published: (2025)
by: Babaei, Hossein, et al.
Published: (2025)
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
by: Tanigawa, Risako, et al.
Published: (2024)
by: Tanigawa, Risako, et al.
Published: (2024)
Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation
by: Premananth, Gowtham, et al.
Published: (2025)
by: Premananth, Gowtham, et al.
Published: (2025)
Similar Items
-
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
by: Yeo, Jeong Hun, et al.
Published: (2023) -
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
by: Salaj, Ina, et al.
Published: (2025) -
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023) -
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024) -
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
by: Mingote, Victoria, et al.
Published: (2024)