Handling Numeric Expressions in Automatic Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huber, Christian, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
von: Min, Do June, et al.
Veröffentlicht: (2024)
von: Min, Do June, et al.
Veröffentlicht: (2024)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
von: Zhang, Shucong, et al.
Veröffentlicht: (2025)
von: Zhang, Shucong, et al.
Veröffentlicht: (2025)
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
von: Le-Duc, Khai
Veröffentlicht: (2024)
von: Le-Duc, Khai
Veröffentlicht: (2024)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
von: Sudarshan, Ankitha, et al.
Veröffentlicht: (2023)
von: Sudarshan, Ankitha, et al.
Veröffentlicht: (2023)
Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
von: S, Chandrashekar M, et al.
Veröffentlicht: (2026)
von: S, Chandrashekar M, et al.
Veröffentlicht: (2026)
Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices
von: Kim, Gwantae, et al.
Veröffentlicht: (2024)
von: Kim, Gwantae, et al.
Veröffentlicht: (2024)
Semantically Corrected Amharic Automatic Speech Recognition
von: Adnew, Samuael, et al.
Veröffentlicht: (2024)
von: Adnew, Samuael, et al.
Veröffentlicht: (2024)
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages
von: Pillai, Leena G, et al.
Veröffentlicht: (2024)
von: Pillai, Leena G, et al.
Veröffentlicht: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling
von: McGuire, Michael
Veröffentlicht: (2025)
von: McGuire, Michael
Veröffentlicht: (2025)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
Towards Robust Speech Recognition for Jamaican Patois Music Transcription
von: Madden, Jordan, et al.
Veröffentlicht: (2025)
von: Madden, Jordan, et al.
Veröffentlicht: (2025)
Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
von: Li, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Li, Xiaoxiao, et al.
Veröffentlicht: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
von: Storey, Edward, et al.
Veröffentlicht: (2025)
von: Storey, Edward, et al.
Veröffentlicht: (2025)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
von: Moure, Pehuén, et al.
Veröffentlicht: (2026)
von: Moure, Pehuén, et al.
Veröffentlicht: (2026)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Luganda Speech Intent Recognition for IoT Applications
von: Katumba, Andrew, et al.
Veröffentlicht: (2024)
von: Katumba, Andrew, et al.
Veröffentlicht: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
Adapting Foundation Speech Recognition Models to Impaired Speech: A Semantic Re-chaining Approach for Personalization of German Speech
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
von: Seide, Frank, et al.
Veröffentlicht: (2024)
von: Seide, Frank, et al.
Veröffentlicht: (2024)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts
von: Gao, Lingyun, et al.
Veröffentlicht: (2025)
von: Gao, Lingyun, et al.
Veröffentlicht: (2025)
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
PhoWhisper: Automatic Speech Recognition for Vietnamese
von: Le, Thanh-Thien, et al.
Veröffentlicht: (2024)
von: Le, Thanh-Thien, et al.
Veröffentlicht: (2024)
Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
von: Wu, Chung-Wen, et al.
Veröffentlicht: (2024)
von: Wu, Chung-Wen, et al.
Veröffentlicht: (2024)
Phonology-Guided Speech-to-Speech Translation for African Languages
von: Ochieng, Peter, et al.
Veröffentlicht: (2024)
von: Ochieng, Peter, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
von: Min, Do June, et al.
Veröffentlicht: (2024) -
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
von: Zhang, Shucong, et al.
Veröffentlicht: (2025) -
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
von: Le-Duc, Khai
Veröffentlicht: (2024) -
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024) -
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
von: Sudarshan, Ankitha, et al.
Veröffentlicht: (2023)