Using Songs to Improve Kazakh Automatic Speech Recognition
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Yeshpanov, Rustem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
Unsupervised Online Continual Learning for Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2024)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Non-Intrusive Automatic Speech Recognition Refinement: A Survey
von: Peyghan, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Peyghan, Mohammad Reza, et al.
Veröffentlicht: (2025)
Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction
von: Folkertsma, Hidde, et al.
Veröffentlicht: (2026)
von: Folkertsma, Hidde, et al.
Veröffentlicht: (2026)
Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2026)
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2026)
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
von: Fu, Li, et al.
Veröffentlicht: (2024)
von: Fu, Li, et al.
Veröffentlicht: (2024)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
Group-Aware Partial Model Merging for Children's Automatic Speech Recognition
von: Rolland, Thomas, et al.
Veröffentlicht: (2025)
von: Rolland, Thomas, et al.
Veröffentlicht: (2025)
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
von: Wang, Pu, et al.
Veröffentlicht: (2024)
von: Wang, Pu, et al.
Veröffentlicht: (2024)
Cross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech Recognition
von: Mitsumori, Shunsuke, et al.
Veröffentlicht: (2025)
von: Mitsumori, Shunsuke, et al.
Veröffentlicht: (2025)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
Automatic Live Music Song Identification Using Multi-level Deep Sequence Similarity Learning
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
Detecting and Defending Against Adversarial Attacks on Automatic Speech Recognition via Diffusion Models
von: Kühne, Nikolai L., et al.
Veröffentlicht: (2024)
von: Kühne, Nikolai L., et al.
Veröffentlicht: (2024)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Rehearsal-Free Online Continual Learning for Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2023)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2023)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
von: Yang, Mu, et al.
Veröffentlicht: (2025)
von: Yang, Mu, et al.
Veröffentlicht: (2025)
Automatic Detection of Depression in Speech Using Ensemble Convolutional Neural Networks
von: Vázquez-Romero, Adrián, et al.
Veröffentlicht: (2024)
von: Vázquez-Romero, Adrián, et al.
Veröffentlicht: (2024)
PhoWhisper: Automatic Speech Recognition for Vietnamese
von: Le, Thanh-Thien, et al.
Veröffentlicht: (2024)
von: Le, Thanh-Thien, et al.
Veröffentlicht: (2024)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
Towards Improved Speech Recognition through Optimized Synthetic Data Generation
von: Perrin, Yanis, et al.
Veröffentlicht: (2025)
von: Perrin, Yanis, et al.
Veröffentlicht: (2025)
SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition for Hindi
von: Saha, Anish, et al.
Veröffentlicht: (2024)
von: Saha, Anish, et al.
Veröffentlicht: (2024)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
von: Chen, Hongzhao, et al.
Veröffentlicht: (2025)
von: Chen, Hongzhao, et al.
Veröffentlicht: (2025)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024) -
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026) -
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
von: Polok, Alexander, et al.
Veröffentlicht: (2024) -
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
von: Li, Jinpeng, et al.
Veröffentlicht: (2024) -
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)