Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Triantafyllopoulos, Andreas, Batliner, Anton, Schuller, Björn W. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
von: Wagner, Philipp, et al.
Veröffentlicht: (2024)
von: Wagner, Philipp, et al.
Veröffentlicht: (2024)
Abusive Speech Detection in Indic Languages Using Acoustic Features
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
An automatic analysis of ultrasound vocalisations for the prediction of interaction context in captive Egyptian fruit bats
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
Computer Audition: From Task-Specific Machine Learning to Foundation Models
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
von: Jing, Xin, et al.
Veröffentlicht: (2026)
von: Jing, Xin, et al.
Veröffentlicht: (2026)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
von: Jing, Xin, et al.
Veröffentlicht: (2025)
von: Jing, Xin, et al.
Veröffentlicht: (2025)
Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
von: Milling, Manuel, et al.
Veröffentlicht: (2024)
von: Milling, Manuel, et al.
Veröffentlicht: (2024)
autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks
von: Rampp, Simon, et al.
Veröffentlicht: (2024)
von: Rampp, Simon, et al.
Veröffentlicht: (2024)
Exploring Meta Information for Audio-based Zero-shot Bird Classification
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
learning discriminative features from spectrograms using center loss for speech emotion recognition
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
Towards interpretable emotion recognition: Identifying key features with machine learning
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
Bringing the Discussion of Minima Sharpness to the Audio Domain: a Filter-Normalised Evaluation for Acoustic Scene Classification
von: Milling, Manuel, et al.
Veröffentlicht: (2023)
von: Milling, Manuel, et al.
Veröffentlicht: (2023)
Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life
von: Batliner, Anton, et al.
Veröffentlicht: (2025)
von: Batliner, Anton, et al.
Veröffentlicht: (2025)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
von: Kashyap, Bipasha, et al.
Veröffentlicht: (2026)
von: Kashyap, Bipasha, et al.
Veröffentlicht: (2026)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
von: Ding, Jiani, et al.
Veröffentlicht: (2025)
von: Ding, Jiani, et al.
Veröffentlicht: (2025)
Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty Estimation
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Can Large Language Models Aid in Annotating Speech Emotional Data? Uncovering New Frontiers
von: Latif, Siddique, et al.
Veröffentlicht: (2023)
von: Latif, Siddique, et al.
Veröffentlicht: (2023)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
Explainable Detection of Machine Generated Music and Early Systematic Evaluation
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
Versatile audio-visual learning for emotion recognition
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
von: Jing, Xin, et al.
Veröffentlicht: (2024) -
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
von: Jing, Xin, et al.
Veröffentlicht: (2024) -
Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
von: Wagner, Philipp, et al.
Veröffentlicht: (2024) -
Abusive Speech Detection in Indic Languages Using Acoustic Features
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024) -
An automatic analysis of ultrasound vocalisations for the prediction of interaction context in captive Egyptian fruit bats
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)