Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dixit, Satvik, Low, Daniel M., Elbanna, Gasser, Catania, Fabio, Ghosh, Satrajit S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Speaker Identity Coding in Self-supervised Models and Humans
von: Elbanna, Gasser
Veröffentlicht: (2024)
von: Elbanna, Gasser
Veröffentlicht: (2024)
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Learning Perceptually Relevant Temporal Envelope Morphing
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
von: Saliba, Alexandra, et al.
Veröffentlicht: (2024)
von: Saliba, Alexandra, et al.
Veröffentlicht: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2025)
On the Contribution of Lexical Features to Speech Emotion Recognition
von: Combei, David
Veröffentlicht: (2025)
von: Combei, David
Veröffentlicht: (2025)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
Abusive Speech Detection in Indic Languages Using Acoustic Features
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
Interpretable Embeddings of Speech Enhance and Explain Brain Encoding Performance of Audio Models
von: Shimizu, Riki, et al.
Veröffentlicht: (2025)
von: Shimizu, Riki, et al.
Veröffentlicht: (2025)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
von: Tan, Chao, et al.
Veröffentlicht: (2024)
von: Tan, Chao, et al.
Veröffentlicht: (2024)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
PCQ: Emotion Recognition in Speech via Progressive Channel Querying
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
von: Derington, Anna, et al.
Veröffentlicht: (2023)
von: Derington, Anna, et al.
Veröffentlicht: (2023)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Acoustic Feature Extraction Tools for Clinical Speech Analysis
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
von: Huang, Mengcheng, et al.
Veröffentlicht: (2026)
von: Huang, Mengcheng, et al.
Veröffentlicht: (2026)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating Speaker Identity Coding in Self-supervised Models and Humans
von: Elbanna, Gasser
Veröffentlicht: (2024) -
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024) -
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024) -
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024) -
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)