Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dhasmana, Akriti, Srivastava, Aarohi, Chiang, David
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917265922850816
author Dhasmana, Akriti
Srivastava, Aarohi
Chiang, David
author_facet Dhasmana, Akriti
Srivastava, Aarohi
Chiang, David
contents We conduct an empirical study of cross-lingual transfer using spontaneous, noisy, and code-mixed speech across a wide range of Indic dialects and language varieties. Our results indicate that although ASR performance is generally improved with reduced phylogenetic distance between languages, this factor alone does not fully explain performance in dialectal settings. Often, fine-tuning on smaller amounts of dialectal data yields performance comparable to fine-tuning on larger amounts of phylogenetically-related, high-resource standardized languages. We also present a case study on Garhwali, a low-resource Pahari language variety, and evaluate multiple contemporary ASR models. Finally, we analyze transcription errors to examine bias toward pre-training languages, providing additional insight into challenges faced by ASR systems on dialectal and non-standardized speech.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04373
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties
Dhasmana, Akriti
Srivastava, Aarohi
Chiang, David
Computation and Language
I.2.7
We conduct an empirical study of cross-lingual transfer using spontaneous, noisy, and code-mixed speech across a wide range of Indic dialects and language varieties. Our results indicate that although ASR performance is generally improved with reduced phylogenetic distance between languages, this factor alone does not fully explain performance in dialectal settings. Often, fine-tuning on smaller amounts of dialectal data yields performance comparable to fine-tuning on larger amounts of phylogenetically-related, high-resource standardized languages. We also present a case study on Garhwali, a low-resource Pahari language variety, and evaluate multiple contemporary ASR models. Finally, we analyze transcription errors to examine bias toward pre-training languages, providing additional insight into challenges faced by ASR systems on dialectal and non-standardized speech.
title Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2601.04373