Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Zhichen, Geng, Tianqi, Feng, Hui, Yuan, Jiahong, Richmond, Korin, Li, Yuanchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
I Know Your Feelings Before You Do: Predicting Future Affective Reactions in Human-Computer Dialogue
von: Li, Yuanchao, et al.
Veröffentlicht: (2023)
von: Li, Yuanchao, et al.
Veröffentlicht: (2023)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
von: Sun, Licai, et al.
Veröffentlicht: (2024)
von: Sun, Licai, et al.
Veröffentlicht: (2024)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
von: Dietrich, Juergen
Veröffentlicht: (2026)
von: Dietrich, Juergen
Veröffentlicht: (2026)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
von: Cai, Zhuojiang, et al.
Veröffentlicht: (2024)
von: Cai, Zhuojiang, et al.
Veröffentlicht: (2024)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
von: Ma, Yong, et al.
Veröffentlicht: (2025)
von: Ma, Yong, et al.
Veröffentlicht: (2025)
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2022)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2022)
Collaboration Between Robots, Interfaces and Humans: Practice-Based and Audience Perspectives
von: Savery, Anna, et al.
Veröffentlicht: (2024)
von: Savery, Anna, et al.
Veröffentlicht: (2024)
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
von: Li, Yinan, et al.
Veröffentlicht: (2026)
von: Li, Yinan, et al.
Veröffentlicht: (2026)
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
von: Saliba, Alexandra, et al.
Veröffentlicht: (2024)
von: Saliba, Alexandra, et al.
Veröffentlicht: (2024)
Are Expressions for Music Emotions the Same Across Cultures?
von: Celen, Elif, et al.
Veröffentlicht: (2025)
von: Celen, Elif, et al.
Veröffentlicht: (2025)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025) -
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024) -
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024) -
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025) -
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)