The evaluation of a code-switched Sepedi-English automatic speech recognition system
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phaladi, Amanda, Modipa, Thipe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025)
von: Linke, Julian, et al.
Veröffentlicht: (2025)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
More than words: Advancements and challenges in speech recognition for singing
von: Kruspe, Anna
Veröffentlicht: (2024)
von: Kruspe, Anna
Veröffentlicht: (2024)
End-to-end streaming model for low-latency speech anonymization
von: Quamer, Waris, et al.
Veröffentlicht: (2024)
von: Quamer, Waris, et al.
Veröffentlicht: (2024)
DarkStream: real-time speech anonymization with low latency
von: Quamer, Waris, et al.
Veröffentlicht: (2025)
von: Quamer, Waris, et al.
Veröffentlicht: (2025)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Speech foundation models in healthcare: Effect of layer selection on pathological speech feature prediction
von: Wiepert, Daniela A., et al.
Veröffentlicht: (2024)
von: Wiepert, Daniela A., et al.
Veröffentlicht: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2024)
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2024)
Pre-training a Transformer-Based Generative Model Using a Small Sepedi Dataset
von: Ramalepe, Simon P., et al.
Veröffentlicht: (2025)
von: Ramalepe, Simon P., et al.
Veröffentlicht: (2025)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
Non-verbal information in spontaneous speech -- towards a new framework of analysis
von: Biron, Tirza, et al.
Veröffentlicht: (2024)
von: Biron, Tirza, et al.
Veröffentlicht: (2024)
Self-supervised learning of speech representations with Dutch archival data
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
von: Fernández-Díaz, Miguel, et al.
Veröffentlicht: (2024)
von: Fernández-Díaz, Miguel, et al.
Veröffentlicht: (2024)
A low latency attention module for streaming self-supervised speech representation learning
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
von: San, Nay, et al.
Veröffentlicht: (2024)
von: San, Nay, et al.
Veröffentlicht: (2024)
Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
von: Jinnai, Yuu
Veröffentlicht: (2025)
von: Jinnai, Yuu
Veröffentlicht: (2025)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
Treble10: A high-quality dataset for far-field speech recognition, dereverberation, and enhancement
von: Mullins, Sarabeth S., et al.
Veröffentlicht: (2025)
von: Mullins, Sarabeth S., et al.
Veröffentlicht: (2025)
Convoifilter: A case study of doing cocktail party speech recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
Perceptual implications of automatic anonymization in pathological speech
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2025)
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2025)
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
How Does a Deep Neural Network Look at Lexical Stress in English Words?
von: Allouche, Itai, et al.
Veröffentlicht: (2025)
von: Allouche, Itai, et al.
Veröffentlicht: (2025)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025) -
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024) -
SeMaScore : a new evaluation metric for automatic speech recognition tasks
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024) -
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023) -
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)