Exploring the limits of decoder-only models trained on public speech recognition corpora
Fuente:
arXiv
Salvato in:
| Autori principali: | Gupta, Ankit, Saon, George, Kingsbury, Brian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Semi-Autoregressive Streaming ASR With Label Context
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
di: Fan, Xulin, et al.
Pubblicazione: (2026)
di: Fan, Xulin, et al.
Pubblicazione: (2026)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
di: Araiza-Illan, Gloria, et al.
Pubblicazione: (2023)
di: Araiza-Illan, Gloria, et al.
Pubblicazione: (2023)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
Improving child speech recognition with augmented child-like speech
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
di: Wright, George August, et al.
Pubblicazione: (2023)
di: Wright, George August, et al.
Pubblicazione: (2023)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
di: Dhakal, Manish, et al.
Pubblicazione: (2024)
di: Dhakal, Manish, et al.
Pubblicazione: (2024)
Self-consistent context aware conformer transducer for speech recognition
di: Kolokolov, Konstantin, et al.
Pubblicazione: (2024)
di: Kolokolov, Konstantin, et al.
Pubblicazione: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
di: Ma, Te, et al.
Pubblicazione: (2025)
di: Ma, Te, et al.
Pubblicazione: (2025)
Convoifilter: A case study of doing cocktail party speech recognition
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2023)
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2023)
Introduction to speech recognition
di: Dauphin, Gabriel
Pubblicazione: (2024)
di: Dauphin, Gabriel
Pubblicazione: (2024)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
di: Sedukhin, Oleg, et al.
Pubblicazione: (2026)
di: Sedukhin, Oleg, et al.
Pubblicazione: (2026)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
di: Jiang, Liuyuan, et al.
Pubblicazione: (2025)
di: Jiang, Liuyuan, et al.
Pubblicazione: (2025)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
di: Banerjee, Sourav, et al.
Pubblicazione: (2024)
di: Banerjee, Sourav, et al.
Pubblicazione: (2024)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
Adjust-free adversarial example generation in speech recognition using evolutionary multi-objective optimization under black-box condition
di: Ishida, Shoma, et al.
Pubblicazione: (2020)
di: Ishida, Shoma, et al.
Pubblicazione: (2020)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
Transferable speech-to-text large language model alignment module
di: Wu, Boyong, et al.
Pubblicazione: (2024)
di: Wu, Boyong, et al.
Pubblicazione: (2024)
ZIPA: A family of efficient models for multilingual phone recognition
di: Zhu, Jian, et al.
Pubblicazione: (2025)
di: Zhu, Jian, et al.
Pubblicazione: (2025)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Exploring the anatomy of articulation rate in spontaneous English speech: relationships between utterance length effects and social factors
di: Tanner, James, et al.
Pubblicazione: (2024)
di: Tanner, James, et al.
Pubblicazione: (2024)
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024)
di: Oneata, Dan, et al.
Pubblicazione: (2024)
AfriHuBERT: A self-supervised speech representation model for African languages
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
di: Sanchez, Ariadna, et al.
Pubblicazione: (2025)
di: Sanchez, Ariadna, et al.
Pubblicazione: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
di: Wan, Xucheng, et al.
Pubblicazione: (2024)
di: Wan, Xucheng, et al.
Pubblicazione: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
di: Lonergan, Liam, et al.
Pubblicazione: (2024)
di: Lonergan, Liam, et al.
Pubblicazione: (2024)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
di: Malisz, Zofia, et al.
Pubblicazione: (2024)
di: Malisz, Zofia, et al.
Pubblicazione: (2024)
Semantic enrichment towards efficient speech representations
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
di: Ling, Shaoshi, et al.
Pubblicazione: (2023)
di: Ling, Shaoshi, et al.
Pubblicazione: (2023)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
di: Liu, Heyang, et al.
Pubblicazione: (2024)
di: Liu, Heyang, et al.
Pubblicazione: (2024)
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024)
di: Kamper, Herman, et al.
Pubblicazione: (2024)
Can Whisper perform speech-based in-context learning?
di: Wang, Siyin, et al.
Pubblicazione: (2023)
di: Wang, Siyin, et al.
Pubblicazione: (2023)
A Non-autoregressive Model for Joint STT and TTS
di: Sunder, Vishal, et al.
Pubblicazione: (2025)
di: Sunder, Vishal, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Semi-Autoregressive Streaming ASR With Label Context
di: Arora, Siddhant, et al.
Pubblicazione: (2023) -
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
di: Fan, Xulin, et al.
Pubblicazione: (2026) -
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
di: Araiza-Illan, Gloria, et al.
Pubblicazione: (2023) -
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023) -
Improving child speech recognition with augmented child-like speech
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)