Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Szu-Jui, Hansen, John H. L. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
learning discriminative features from spectrograms using center loss for speech emotion recognition
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Towards interpretable emotion recognition: Identifying key features with machine learning
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
von: Samui, Suman
Veröffentlicht: (2022)
von: Samui, Suman
Veröffentlicht: (2022)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Dementia classification from spontaneous speech using wrapper-based feature selection
von: Niemelä, Marko, et al.
Veröffentlicht: (2025)
von: Niemelä, Marko, et al.
Veröffentlicht: (2025)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
AxLSTMs: learning self-supervised audio representations with xLSTMs
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
Distilling a speech and music encoder with task arithmetic
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Forensic deepfake audio detection using segmental speech features
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021) -
learning discriminative features from spectrograms using center loss for speech emotion recognition
von: Dai, Dongyang, et al.
Veröffentlicht: (2025) -
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025) -
Towards interpretable emotion recognition: Identifying key features with machine learning
von: Kaloga, Yacouba, et al.
Veröffentlicht: (2025) -
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)