Introduction to speech recognition
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Dauphin, Gabriel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
More than words: Advancements and challenges in speech recognition for singing
von: Kruspe, Anna
Veröffentlicht: (2024)
von: Kruspe, Anna
Veröffentlicht: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Self-supervised learning of speech representations with Dutch archival data
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
Convoifilter: A case study of doing cocktail party speech recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
Non-verbal information in spontaneous speech -- towards a new framework of analysis
von: Biron, Tirza, et al.
Veröffentlicht: (2024)
von: Biron, Tirza, et al.
Veröffentlicht: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
A low latency attention module for streaming self-supervised speech representation learning
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
von: Dhakal, Manish, et al.
Veröffentlicht: (2024)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
Adjust-free adversarial example generation in speech recognition using evolutionary multi-objective optimization under black-box condition
von: Ishida, Shoma, et al.
Veröffentlicht: (2020)
von: Ishida, Shoma, et al.
Veröffentlicht: (2020)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
Technical Report on classification of literature related to children speech disorder
von: Wang, Ziang, et al.
Veröffentlicht: (2025)
von: Wang, Ziang, et al.
Veröffentlicht: (2025)
Moshi: a speech-text foundation model for real-time dialogue
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
Towards measuring fairness in speech recognition: Fair-Speech dataset
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
Translating speech with just images
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
Textually Pretrained Speech Language Models
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
von: Zhang, Kevin, et al.
Veröffentlicht: (2024)
von: Zhang, Kevin, et al.
Veröffentlicht: (2024)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2024)
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024) -
More than words: Advancements and challenges in speech recognition for singing
von: Kruspe, Anna
Veröffentlicht: (2024) -
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024) -
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024) -
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)