Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911238844317696 |
|---|---|
| author | Dipto, Tawsif Tashwar Hossain, Azmol Faruque, Rubayet Sabbir Hassan, Md. Rezuwan Fatema, Kanij Shome, Tanmoy Naswan, Ruwad Zihad, Md. Foriduzzaman Anam, Mohaymen Ul Tasnim, Nazia Mahmud, Hasan Hasan, Md Kamrul Shawon, Md. Mehedi Hasan Sadeque, Farig Reasat, Tahsin |
| author_facet | Dipto, Tawsif Tashwar Hossain, Azmol Faruque, Rubayet Sabbir Hassan, Md. Rezuwan Fatema, Kanij Shome, Tanmoy Naswan, Ruwad Zihad, Md. Foriduzzaman Anam, Mohaymen Ul Tasnim, Nazia Mahmud, Hasan Hasan, Md Kamrul Shawon, Md. Mehedi Hasan Sadeque, Farig Reasat, Tahsin |
| contents | Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is treated as a fine-tuning task. To investigate the effects of dialectal variations on ASR we develop a 78-hour annotated Bengali Speech-to-Text (STT) corpus named Ben-10. Investigation from linguistic and data-driven perspectives shows that speech foundation models struggle heavily in regional dialect ASR, both in zero-shot and fine-tuned settings. We observe that all deep learning methods struggle to model speech data under dialectal variations but dialect specific model training alleviates the issue. Our dataset also serves as a out of-distribution (OOD) resource for ASR modeling under constrained resources in ASR algorithms. The dataset and code developed for this project are publicly available |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_23252 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages? Dipto, Tawsif Tashwar Hossain, Azmol Faruque, Rubayet Sabbir Hassan, Md. Rezuwan Fatema, Kanij Shome, Tanmoy Naswan, Ruwad Zihad, Md. Foriduzzaman Anam, Mohaymen Ul Tasnim, Nazia Mahmud, Hasan Hasan, Md Kamrul Shawon, Md. Mehedi Hasan Sadeque, Farig Reasat, Tahsin Computation and Language Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is treated as a fine-tuning task. To investigate the effects of dialectal variations on ASR we develop a 78-hour annotated Bengali Speech-to-Text (STT) corpus named Ben-10. Investigation from linguistic and data-driven perspectives shows that speech foundation models struggle heavily in regional dialect ASR, both in zero-shot and fine-tuned settings. We observe that all deep learning methods struggle to model speech data under dialectal variations but dialect specific model training alleviates the issue. Our dataset also serves as a out of-distribution (OOD) resource for ASR modeling under constrained resources in ASR algorithms. The dataset and code developed for this project are publicly available |
| title | Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages? |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.23252 |