Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dipto, Tawsif Tashwar, Hossain, Azmol, Faruque, Rubayet Sabbir, Hassan, Md. Rezuwan, Fatema, Kanij, Shome, Tanmoy, Naswan, Ruwad, Zihad, Md. Foriduzzaman, Anam, Mohaymen Ul, Tasnim, Nazia, Mahmud, Hasan, Hasan, Md Kamrul, Shawon, Md. Mehedi Hasan, Sadeque, Farig, Reasat, Tahsin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911238844317696
author Dipto, Tawsif Tashwar
Hossain, Azmol
Faruque, Rubayet Sabbir
Hassan, Md. Rezuwan
Fatema, Kanij
Shome, Tanmoy
Naswan, Ruwad
Zihad, Md. Foriduzzaman
Anam, Mohaymen Ul
Tasnim, Nazia
Mahmud, Hasan
Hasan, Md Kamrul
Shawon, Md. Mehedi Hasan
Sadeque, Farig
Reasat, Tahsin
author_facet Dipto, Tawsif Tashwar
Hossain, Azmol
Faruque, Rubayet Sabbir
Hassan, Md. Rezuwan
Fatema, Kanij
Shome, Tanmoy
Naswan, Ruwad
Zihad, Md. Foriduzzaman
Anam, Mohaymen Ul
Tasnim, Nazia
Mahmud, Hasan
Hasan, Md Kamrul
Shawon, Md. Mehedi Hasan
Sadeque, Farig
Reasat, Tahsin
contents Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is treated as a fine-tuning task. To investigate the effects of dialectal variations on ASR we develop a 78-hour annotated Bengali Speech-to-Text (STT) corpus named Ben-10. Investigation from linguistic and data-driven perspectives shows that speech foundation models struggle heavily in regional dialect ASR, both in zero-shot and fine-tuned settings. We observe that all deep learning methods struggle to model speech data under dialectal variations but dialect specific model training alleviates the issue. Our dataset also serves as a out of-distribution (OOD) resource for ASR modeling under constrained resources in ASR algorithms. The dataset and code developed for this project are publicly available
format Preprint
id arxiv_https___arxiv_org_abs_2510_23252
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
Dipto, Tawsif Tashwar
Hossain, Azmol
Faruque, Rubayet Sabbir
Hassan, Md. Rezuwan
Fatema, Kanij
Shome, Tanmoy
Naswan, Ruwad
Zihad, Md. Foriduzzaman
Anam, Mohaymen Ul
Tasnim, Nazia
Mahmud, Hasan
Hasan, Md Kamrul
Shawon, Md. Mehedi Hasan
Sadeque, Farig
Reasat, Tahsin
Computation and Language
Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is treated as a fine-tuning task. To investigate the effects of dialectal variations on ASR we develop a 78-hour annotated Bengali Speech-to-Text (STT) corpus named Ben-10. Investigation from linguistic and data-driven perspectives shows that speech foundation models struggle heavily in regional dialect ASR, both in zero-shot and fine-tuned settings. We observe that all deep learning methods struggle to model speech data under dialectal variations but dialect specific model training alleviates the issue. Our dataset also serves as a out of-distribution (OOD) resource for ASR modeling under constrained resources in ASR algorithms. The dataset and code developed for this project are publicly available
title Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
topic Computation and Language
url https://arxiv.org/abs/2510.23252