Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rustagi, Arnav, Bajpai, Satvik, Kaur, Nimrat, Siddharth, Siddharth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915320166350848
author Rustagi, Arnav
Bajpai, Satvik
Kaur, Nimrat
Siddharth, Siddharth
author_facet Rustagi, Arnav
Bajpai, Satvik
Kaur, Nimrat
Siddharth, Siddharth
contents Computer-Assisted Pronunciation Training (CAPT) has been extensively studied for English. However, there remains a critical gap in its application to Indian languages with a base of 1.5 billion speakers. Pronunciation tools tailored to Indian languages are strikingly lacking despite the fact that millions learn them every year. With over 600 million speakers and being the fourth most-spoken language worldwide, improving Hindi pronunciation is a vital first step toward addressing this gap. This paper proposes 1) Dhvani -- a novel CAPT system for Hindi, 2) synthetic speech generation for Hindi mispronunciations, and 3) a novel methodology for providing personalized feedback to learners. While the system often interacts with learners using Devanagari graphemes, its core analysis targets phonemic distinctions, leveraging Hindi's highly phonetic orthography to analyze mispronounced speech and provide targeted feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02166
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi
Rustagi, Arnav
Bajpai, Satvik
Kaur, Nimrat
Siddharth, Siddharth
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Computer-Assisted Pronunciation Training (CAPT) has been extensively studied for English. However, there remains a critical gap in its application to Indian languages with a base of 1.5 billion speakers. Pronunciation tools tailored to Indian languages are strikingly lacking despite the fact that millions learn them every year. With over 600 million speakers and being the fourth most-spoken language worldwide, improving Hindi pronunciation is a vital first step toward addressing this gap. This paper proposes 1) Dhvani -- a novel CAPT system for Hindi, 2) synthetic speech generation for Hindi mispronunciations, and 3) a novel methodology for providing personalized feedback to learners. While the system often interacts with learners using Devanagari graphemes, its core analysis targets phonemic distinctions, leveraging Hindi's highly phonetic orthography to analyze mispronounced speech and provide targeted feedback.
title Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.02166