Self-consistent context aware conformer transducer for speech recognition
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kolokolov, Konstantin, Pekichev, Pavel, Raghunathan, Karthik |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
par: Pandey, Rahul, et autres
Publié: (2023)
par: Pandey, Rahul, et autres
Publié: (2023)
Improving child speech recognition with augmented child-like speech
par: Zhang, Yuanyuan, et autres
Publié: (2024)
par: Zhang, Yuanyuan, et autres
Publié: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
par: Zheng, Naijun, et autres
Publié: (2024)
par: Zheng, Naijun, et autres
Publié: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
par: Dong, Lukuang, et autres
Publié: (2026)
par: Dong, Lukuang, et autres
Publié: (2026)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
par: Ma, Te, et autres
Publié: (2025)
par: Ma, Te, et autres
Publié: (2025)
Convoifilter: A case study of doing cocktail party speech recognition
par: Nguyen, Thai-Binh, et autres
Publié: (2023)
par: Nguyen, Thai-Binh, et autres
Publié: (2023)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
par: Araiza-Illan, Gloria, et autres
Publié: (2023)
par: Araiza-Illan, Gloria, et autres
Publié: (2023)
Can Whisper perform speech-based in-context learning?
par: Wang, Siyin, et autres
Publié: (2023)
par: Wang, Siyin, et autres
Publié: (2023)
Introduction to speech recognition
par: Dauphin, Gabriel
Publié: (2024)
par: Dauphin, Gabriel
Publié: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
par: Gupta, Ankit, et autres
Publié: (2024)
par: Gupta, Ankit, et autres
Publié: (2024)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
par: Sedukhin, Oleg, et autres
Publié: (2026)
par: Sedukhin, Oleg, et autres
Publié: (2026)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
par: Dhakal, Manish, et autres
Publié: (2024)
par: Dhakal, Manish, et autres
Publié: (2024)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
par: Banerjee, Sourav, et autres
Publié: (2024)
par: Banerjee, Sourav, et autres
Publié: (2024)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
par: Wright, George August, et autres
Publié: (2023)
par: Wright, George August, et autres
Publié: (2023)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
par: Wang, Hsuan-Fu, et autres
Publié: (2024)
par: Wang, Hsuan-Fu, et autres
Publié: (2024)
Adjust-free adversarial example generation in speech recognition using evolutionary multi-objective optimization under black-box condition
par: Ishida, Shoma, et autres
Publié: (2020)
par: Ishida, Shoma, et autres
Publié: (2020)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
par: Liu, Oli Danyi, et autres
Publié: (2024)
par: Liu, Oli Danyi, et autres
Publié: (2024)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
par: Gowda, Harshavardhana T., et autres
Publié: (2025)
par: Gowda, Harshavardhana T., et autres
Publié: (2025)
Translating speech with just images
par: Oneata, Dan, et autres
Publié: (2024)
par: Oneata, Dan, et autres
Publié: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
par: Okocha, Chibuzor, et autres
Publié: (2025)
par: Okocha, Chibuzor, et autres
Publié: (2025)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
par: Wan, Xucheng, et autres
Publié: (2024)
par: Wan, Xucheng, et autres
Publié: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
par: Lonergan, Liam, et autres
Publié: (2024)
par: Lonergan, Liam, et autres
Publié: (2024)
Semantic enrichment towards efficient speech representations
par: Laperrière, Gaëlle, et autres
Publié: (2023)
par: Laperrière, Gaëlle, et autres
Publié: (2023)
Revisiting speech segmentation and lexicon learning with better features
par: Kamper, Herman, et autres
Publié: (2024)
par: Kamper, Herman, et autres
Publié: (2024)
Prominence-aware automatic speech recognition for conversational speech
par: Linke, Julian, et autres
Publié: (2025)
par: Linke, Julian, et autres
Publié: (2025)
Direct Punjabi to English speech translation using discrete units
par: Kaur, Prabhjot, et autres
Publié: (2024)
par: Kaur, Prabhjot, et autres
Publié: (2024)
Transferable speech-to-text large language model alignment module
par: Wu, Boyong, et autres
Publié: (2024)
par: Wu, Boyong, et autres
Publié: (2024)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
par: Lyth, Dan, et autres
Publié: (2024)
par: Lyth, Dan, et autres
Publié: (2024)
Beyond the binary: Limitations and possibilities of gender-related speech technology research
par: Sanchez, Ariadna, et autres
Publié: (2024)
par: Sanchez, Ariadna, et autres
Publié: (2024)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
par: Cui, Ziyun, et autres
Publié: (2023)
par: Cui, Ziyun, et autres
Publié: (2023)
An experiment on an automated literature survey of data-driven speech enhancement methods
par: Santos, Arthur dos, et autres
Publié: (2023)
par: Santos, Arthur dos, et autres
Publié: (2023)
ZIPA: A family of efficient models for multilingual phone recognition
par: Zhu, Jian, et autres
Publié: (2025)
par: Zhu, Jian, et autres
Publié: (2025)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
par: Garg, Abhinav, et autres
Publié: (2024)
par: Garg, Abhinav, et autres
Publié: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
par: Slabbert, Danel, et autres
Publié: (2025)
par: Slabbert, Danel, et autres
Publié: (2025)
AfriHuBERT: A self-supervised speech representation model for African languages
par: Alabi, Jesujoba O., et autres
Publié: (2024)
par: Alabi, Jesujoba O., et autres
Publié: (2024)
Korean aegyo speech shows systematic F1 increase to signal childlike qualities
par: Kim, Ji-eun, et autres
Publié: (2026)
par: Kim, Ji-eun, et autres
Publié: (2026)
Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?
par: Törö, Tuukka, et autres
Publié: (2025)
par: Törö, Tuukka, et autres
Publié: (2025)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
par: Sanchez, Ariadna, et autres
Publié: (2025)
par: Sanchez, Ariadna, et autres
Publié: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
par: Ducorroy, Alexandre, et autres
Publié: (2025)
par: Ducorroy, Alexandre, et autres
Publié: (2025)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
par: Paraskevopoulos, Georgios, et autres
Publié: (2024)
par: Paraskevopoulos, Georgios, et autres
Publié: (2024)
Documents similaires
-
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
par: Pandey, Rahul, et autres
Publié: (2023) -
Improving child speech recognition with augmented child-like speech
par: Zhang, Yuanyuan, et autres
Publié: (2024) -
An efficient text augmentation approach for contextualized Mandarin speech recognition
par: Zheng, Naijun, et autres
Publié: (2024) -
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
par: Dong, Lukuang, et autres
Publié: (2026) -
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
par: Ma, Te, et autres
Publié: (2025)