Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhogale, Kaushal Santosh, Javed, Tahir, John, Greeshma Susan, Rathi, Dhruv, Padmanaban, Akshayasree, Parasa, Niharika, Khapra, Mitesh M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
di: Bhogale, Kaushal Santosh, et al.
Pubblicazione: (2024)
di: Bhogale, Kaushal Santosh, et al.
Pubblicazione: (2024)
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
di: Javed, Tahir, et al.
Pubblicazione: (2025)
di: Javed, Tahir, et al.
Pubblicazione: (2025)
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
di: Bhogale, Kaushal, et al.
Pubblicazione: (2026)
di: Bhogale, Kaushal, et al.
Pubblicazione: (2026)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
di: Joshi, Sakshi, et al.
Pubblicazione: (2025)
di: Joshi, Sakshi, et al.
Pubblicazione: (2025)
LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR Systems
di: Javed, Tahir, et al.
Pubblicazione: (2024)
di: Javed, Tahir, et al.
Pubblicazione: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2025)
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2025)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
di: Anand, Srija, et al.
Pubblicazione: (2024)
di: Anand, Srija, et al.
Pubblicazione: (2024)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
di: Sankar, Ashwin, et al.
Pubblicazione: (2024)
di: Sankar, Ashwin, et al.
Pubblicazione: (2024)
Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
di: Ren, Zhao, et al.
Pubblicazione: (2025)
di: Ren, Zhao, et al.
Pubblicazione: (2025)
Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
di: Kim, June-Woo, et al.
Pubblicazione: (2026)
di: Kim, June-Woo, et al.
Pubblicazione: (2026)
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
di: Kumar, Subham, et al.
Pubblicazione: (2025)
di: Kumar, Subham, et al.
Pubblicazione: (2025)
The State Of TTS: A Case Study with Human Fooling Rates
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2025)
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2025)
Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
di: Javed, Tahir, et al.
Pubblicazione: (2024)
di: Javed, Tahir, et al.
Pubblicazione: (2024)
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
di: Sun, Haoqin, et al.
Pubblicazione: (2026)
di: Sun, Haoqin, et al.
Pubblicazione: (2026)
LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
di: Yoon, Eunseop, et al.
Pubblicazione: (2024)
di: Yoon, Eunseop, et al.
Pubblicazione: (2024)
Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer
di: Jin, Weifei, et al.
Pubblicazione: (2024)
di: Jin, Weifei, et al.
Pubblicazione: (2024)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
di: Park, ChaeHun, et al.
Pubblicazione: (2024)
di: Park, ChaeHun, et al.
Pubblicazione: (2024)
Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
di: Wiafe, Isaac, et al.
Pubblicazione: (2026)
di: Wiafe, Isaac, et al.
Pubblicazione: (2026)
From Hype to Insight: Rethinking Large Language Model Integration in Visual Speech Recognition
di: Jain, Rishabh, et al.
Pubblicazione: (2025)
di: Jain, Rishabh, et al.
Pubblicazione: (2025)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2026)
di: Chen, Youjun, et al.
Pubblicazione: (2026)
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
di: Huang, Dawei, et al.
Pubblicazione: (2026)
di: Huang, Dawei, et al.
Pubblicazione: (2026)
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
di: Lin, Ju, et al.
Pubblicazione: (2024)
di: Lin, Ju, et al.
Pubblicazione: (2024)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
di: Wang, Peng, et al.
Pubblicazione: (2026)
di: Wang, Peng, et al.
Pubblicazione: (2026)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
di: Zink, Oswald, et al.
Pubblicazione: (2024)
di: Zink, Oswald, et al.
Pubblicazione: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
Selective Masking Adversarial Attack on Automatic Speech Recognition Systems
di: Fang, Zheng, et al.
Pubblicazione: (2025)
di: Fang, Zheng, et al.
Pubblicazione: (2025)
Physics-Informed Neural Networks for Speech Production
di: Yokota, Kazuya, et al.
Pubblicazione: (2025)
di: Yokota, Kazuya, et al.
Pubblicazione: (2025)
EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
di: Jing, Xin, et al.
Pubblicazione: (2026)
di: Jing, Xin, et al.
Pubblicazione: (2026)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
Towards Unsupervised Speech Recognition Without Pronunciation Models
di: Ni, Junrui, et al.
Pubblicazione: (2024)
di: Ni, Junrui, et al.
Pubblicazione: (2024)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
Direct Speech to Speech Translation: A Review
di: Sarim, Mohammad, et al.
Pubblicazione: (2025)
di: Sarim, Mohammad, et al.
Pubblicazione: (2025)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2024)
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2024)
In-Materia Speech Recognition
di: Zolfagharinejad, Mohamadreza, et al.
Pubblicazione: (2024)
di: Zolfagharinejad, Mohamadreza, et al.
Pubblicazione: (2024)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
di: Baumann, Ilja, et al.
Pubblicazione: (2025)
di: Baumann, Ilja, et al.
Pubblicazione: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
di: Bhogale, Kaushal Santosh, et al.
Pubblicazione: (2024) -
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
di: Javed, Tahir, et al.
Pubblicazione: (2025) -
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
di: Bhogale, Kaushal, et al.
Pubblicazione: (2026) -
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024) -
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
di: Joshi, Sakshi, et al.
Pubblicazione: (2025)