Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Saliba, Alexandra, Li, Yuanchao, Sanabria, Ramon, Lai, Catherine |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Speech Emotion Recognition with ASR Integration
por: Li, Yuanchao
Publicado: (2026)
por: Li, Yuanchao
Publicado: (2026)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
por: Sun, Yujia, et al.
Publicado: (2024)
por: Sun, Yujia, et al.
Publicado: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
por: Meghanani, Amit, et al.
Publicado: (2024)
por: Meghanani, Amit, et al.
Publicado: (2024)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
por: Ashihara, Takanori, et al.
Publicado: (2024)
por: Ashihara, Takanori, et al.
Publicado: (2024)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
por: Wang, Haoyu, et al.
Publicado: (2022)
por: Wang, Haoyu, et al.
Publicado: (2022)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
por: Han, Zhichen, et al.
Publicado: (2024)
por: Han, Zhichen, et al.
Publicado: (2024)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
por: Hwang, Min-Jae, et al.
Publicado: (2024)
por: Hwang, Min-Jae, et al.
Publicado: (2024)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
por: Wang, Yujin, et al.
Publicado: (2022)
por: Wang, Yujin, et al.
Publicado: (2022)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
por: Zhao, Ya, et al.
Publicado: (2026)
por: Zhao, Ya, et al.
Publicado: (2026)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
por: Xu, Tianyi, et al.
Publicado: (2025)
por: Xu, Tianyi, et al.
Publicado: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
por: Park, Chanho, et al.
Publicado: (2023)
por: Park, Chanho, et al.
Publicado: (2023)
A Cross-Corpus Speech Emotion Recognition Method Based on Supervised Contrastive Learning
por: minjie, Xiang
Publicado: (2024)
por: minjie, Xiang
Publicado: (2024)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
por: Zhang, Enshi, et al.
Publicado: (2024)
por: Zhang, Enshi, et al.
Publicado: (2024)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
por: Shi, Xiaohan, et al.
Publicado: (2023)
por: Shi, Xiaohan, et al.
Publicado: (2023)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
por: Ueda, Lucas H., et al.
Publicado: (2026)
por: Ueda, Lucas H., et al.
Publicado: (2026)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
por: Jiang, Liuyuan, et al.
Publicado: (2025)
por: Jiang, Liuyuan, et al.
Publicado: (2025)
On the Contribution of Lexical Features to Speech Emotion Recognition
por: Combei, David
Publicado: (2025)
por: Combei, David
Publicado: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
por: Hu, Ke, et al.
Publicado: (2025)
por: Hu, Ke, et al.
Publicado: (2025)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
por: Park, Chanho, et al.
Publicado: (2024)
por: Park, Chanho, et al.
Publicado: (2024)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
por: Khaertdinov, Bulat, et al.
Publicado: (2024)
por: Khaertdinov, Bulat, et al.
Publicado: (2024)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
por: Zhang, Zixing, et al.
Publicado: (2024)
por: Zhang, Zixing, et al.
Publicado: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
por: Yamashita, Natsuo, et al.
Publicado: (2024)
por: Yamashita, Natsuo, et al.
Publicado: (2024)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
por: Siriwardhana, Shamane, et al.
Publicado: (2020)
por: Siriwardhana, Shamane, et al.
Publicado: (2020)
Is Self-Supervised Learning Enough to Fill in the Gap? A Study on Speech Inpainting
por: Asaad, Ihab, et al.
Publicado: (2024)
por: Asaad, Ihab, et al.
Publicado: (2024)
Interface Design for Self-Supervised Speech Models
por: Shih, Yi-Jen, et al.
Publicado: (2024)
por: Shih, Yi-Jen, et al.
Publicado: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
por: Lin, Hsi-Che, et al.
Publicado: (2024)
por: Lin, Hsi-Che, et al.
Publicado: (2024)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
por: Nozawa, Kento, et al.
Publicado: (2024)
por: Nozawa, Kento, et al.
Publicado: (2024)
Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
por: Chen, Weidong, et al.
Publicado: (2023)
por: Chen, Weidong, et al.
Publicado: (2023)
BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
por: Tuttösí, Paige, et al.
Publicado: (2025)
por: Tuttösí, Paige, et al.
Publicado: (2025)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
por: Shu, Yuchun, et al.
Publicado: (2024)
por: Shu, Yuchun, et al.
Publicado: (2024)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
por: Venkateswaran, Nitin, et al.
Publicado: (2025)
por: Venkateswaran, Nitin, et al.
Publicado: (2025)
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
por: Yang, Zhengdong, et al.
Publicado: (2025)
por: Yang, Zhengdong, et al.
Publicado: (2025)
Ejemplares similares
-
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
por: Li, Yuanchao, et al.
Publicado: (2024) -
Speech Emotion Recognition with ASR Integration
por: Li, Yuanchao
Publicado: (2026) -
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
por: Sun, Yujia, et al.
Publicado: (2024) -
Crossmodal ASR Error Correction with Discrete Speech Units
por: Li, Yuanchao, et al.
Publicado: (2024) -
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
por: Meghanani, Amit, et al.
Publicado: (2024)