learning discriminative features from spectrograms using center loss for speech emotion recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dai, Dongyang, Wu, Zhiyong, Li, Runnan, Wu, Xixin, Jia, Jia, Meng, Helen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
von: Liu, Peng, et al.
Veröffentlicht: (2024)
von: Liu, Peng, et al.
Veröffentlicht: (2024)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
Towards interpretable emotion recognition: Identifying key features with machine learning
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer
von: Kim, Jeong-Yoon, et al.
Veröffentlicht: (2024)
von: Kim, Jeong-Yoon, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Forensic deepfake audio detection using segmental speech features
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
SongCreator: Lyrics-based Universal Song Generation
von: Lei, Shun, et al.
Veröffentlicht: (2024)
von: Lei, Shun, et al.
Veröffentlicht: (2024)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation
von: Lin, Zhiwei, et al.
Veröffentlicht: (2024)
von: Lin, Zhiwei, et al.
Veröffentlicht: (2024)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
Ensemble of classifiers for speech evaluation
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms
von: Wang, Heehwan, et al.
Veröffentlicht: (2024)
von: Wang, Heehwan, et al.
Veröffentlicht: (2024)
UniSep: Universal Target Audio Separation with Language Models at Scale
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
Automated evaluation of children's speech fluency for low-resource languages
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Selective Classifier-free Guidance for Zero-shot Text-to-speech
von: Zheng, John, et al.
Veröffentlicht: (2025)
von: Zheng, John, et al.
Veröffentlicht: (2025)
FINALLY: fast and universal speech enhancement with studio-like quality
von: Babaev, Nicholas, et al.
Veröffentlicht: (2024)
von: Babaev, Nicholas, et al.
Veröffentlicht: (2024)
Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2025)
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
von: Dai, Dongyang, et al.
Veröffentlicht: (2025) -
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025) -
RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
von: Liu, Peng, et al.
Veröffentlicht: (2024) -
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024) -
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)