LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Javed, Tahir, Nawale, Janki, Joshi, Sakshi, George, Eldho, Bhogale, Kaushal, Mehendale, Deovrat, Khapra, Mitesh M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
by: Joshi, Sakshi, et al.
Published: (2025)
by: Joshi, Sakshi, et al.
Published: (2025)
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
by: Javed, Tahir, et al.
Published: (2025)
by: Javed, Tahir, et al.
Published: (2025)
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
by: Bhogale, Kaushal Santosh, et al.
Published: (2024)
by: Bhogale, Kaushal Santosh, et al.
Published: (2024)
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
by: Bhogale, Kaushal Santosh, et al.
Published: (2026)
by: Bhogale, Kaushal Santosh, et al.
Published: (2026)
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
by: Javed, Tahir, et al.
Published: (2024)
by: Javed, Tahir, et al.
Published: (2024)
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
by: Nawale, Janki Atul, et al.
Published: (2025)
by: Nawale, Janki Atul, et al.
Published: (2025)
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
by: Bhogale, Kaushal, et al.
Published: (2026)
by: Bhogale, Kaushal, et al.
Published: (2026)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
by: Sankar, Ashwin, et al.
Published: (2024)
by: Sankar, Ashwin, et al.
Published: (2024)
Airavata: Introducing Hindi Instruction-tuned LLM
by: Gala, Jay, et al.
Published: (2024)
by: Gala, Jay, et al.
Published: (2024)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
Can Vision-Language Models Evaluate Handwritten Math?
by: Nath, Oikantik, et al.
Published: (2025)
by: Nath, Oikantik, et al.
Published: (2025)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
by: Singh, Anushka, et al.
Published: (2024)
by: Singh, Anushka, et al.
Published: (2024)
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
by: Anand, Srija, et al.
Published: (2024)
by: Anand, Srija, et al.
Published: (2024)
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
by: Anand, Srija, et al.
Published: (2024)
by: Anand, Srija, et al.
Published: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
by: Acharya, Arkadeep, et al.
Published: (2024)
by: Acharya, Arkadeep, et al.
Published: (2024)
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
by: Sankar, Ashwin, et al.
Published: (2025)
by: Sankar, Ashwin, et al.
Published: (2025)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
by: Kamath, Anusha, et al.
Published: (2025)
by: Kamath, Anusha, et al.
Published: (2025)
An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models
by: Mundra, Nandini, et al.
Published: (2024)
by: Mundra, Nandini, et al.
Published: (2024)
The State Of TTS: A Case Study with Human Fooling Rates
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
by: Acharya, Arkadeep, et al.
Published: (2024)
by: Acharya, Arkadeep, et al.
Published: (2024)
HindiLLM: Large Language Model for Hindi
by: Chouhan, Sanjay, et al.
Published: (2024)
by: Chouhan, Sanjay, et al.
Published: (2024)
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
by: Sedghiyeh, Nima, et al.
Published: (2025)
by: Sedghiyeh, Nima, et al.
Published: (2025)
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
by: Sankar, Ashwin, et al.
Published: (2024)
by: Sankar, Ashwin, et al.
Published: (2024)
Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
by: Maurya, Kaushal Kumar, et al.
Published: (2025)
by: Maurya, Kaushal Kumar, et al.
Published: (2025)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
by: Wei, Victor Junqiu, et al.
Published: (2024)
by: Wei, Victor Junqiu, et al.
Published: (2024)
Intonation between phrasing and accent
by: Buchholz, Timo
Published: (2024)
by: Buchholz, Timo
Published: (2024)
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
Multi-class Regret Detection in Hindi Devanagari Script
by: Sharma, Renuka, et al.
Published: (2024)
by: Sharma, Renuka, et al.
Published: (2024)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
by: Nath, Oikantik, et al.
Published: (2025)
by: Nath, Oikantik, et al.
Published: (2025)
On the Robust Approximation of ASR Metrics
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
'Since Lawyers are Males..': Examining Implicit Gender Bias in Hindi Language Generation by LLMs
by: Joshi, Ishika, et al.
Published: (2024)
by: Joshi, Ishika, et al.
Published: (2024)
A Benchmark of French ASR Systems Based on Error Severity
by: Tholly, Antoine, et al.
Published: (2025)
by: Tholly, Antoine, et al.
Published: (2025)
Efficacy of Large Language Models in Systematic Reviews
by: Shah, Aaditya, et al.
Published: (2024)
by: Shah, Aaditya, et al.
Published: (2024)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
by: Gupta, Ashray, et al.
Published: (2025)
by: Gupta, Ashray, et al.
Published: (2025)
Investigating Transcription Normalization in the Faetar ASR Benchmark
by: Peckham, Leo, et al.
Published: (2025)
by: Peckham, Leo, et al.
Published: (2025)
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
Similar Items
-
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
by: Joshi, Sakshi, et al.
Published: (2025) -
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
by: Javed, Tahir, et al.
Published: (2025) -
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
by: Bhogale, Kaushal Santosh, et al.
Published: (2024) -
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
by: Bhogale, Kaushal Santosh, et al.
Published: (2026) -
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
by: Javed, Tahir, et al.
Published: (2024)