Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Tiantian, Huang, Kevin, Xu, Anfeng, Shi, Xuan, Lertpetchpun, Thanathai, Lee, Jihwan, Lee, Yoonjeong, Byrd, Dani, Narayanan, Shrikanth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912518074531840
author Feng, Tiantian
Huang, Kevin
Xu, Anfeng
Shi, Xuan
Lertpetchpun, Thanathai
Lee, Jihwan
Lee, Yoonjeong
Byrd, Dani
Narayanan, Shrikanth
author_facet Feng, Tiantian
Huang, Kevin
Xu, Anfeng
Shi, Xuan
Lertpetchpun, Thanathai
Lee, Jihwan
Lee, Yoonjeong
Byrd, Dani
Narayanan, Shrikanth
contents We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the license of the RAIL family at: https://github.com/tiantiaf0627/voxlect.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01691
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
Feng, Tiantian
Huang, Kevin
Xu, Anfeng
Shi, Xuan
Lertpetchpun, Thanathai
Lee, Jihwan
Lee, Yoonjeong
Byrd, Dani
Narayanan, Shrikanth
Sound
Computation and Language
Audio and Speech Processing
We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the license of the RAIL family at: https://github.com/tiantiaf0627/voxlect.
title Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2508.01691