Towards Unsupervised Speech Recognition at the Syllable-Level

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Liming, Ni, Junrui, Chang, Kai-Wei, Bhati, Saurabhchand, Harwath, David, Hasegawa-Johnson, Mark, Glass, James R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911191776886784
author Wang, Liming
Ni, Junrui
Chang, Kai-Wei
Bhati, Saurabhchand
Harwath, David
Hasegawa-Johnson, Mark
Glass, James R.
author_facet Wang, Liming
Ni, Junrui
Chang, Kai-Wei
Bhati, Saurabhchand
Harwath, David
Hasegawa-Johnson, Mark
Glass, James R.
contents Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning from non-parallel data. However, existing approaches based on phones often rely on costly resources such as grapheme-to-phoneme converters (G2Ps) and struggle to generalize to languages with ambiguous phoneme boundaries due to training instability. In this paper, we address both challenges by introducing a syllable-level UASR framework based on masked language modeling, which avoids the need for G2P and the instability of GAN-based methods. Our approach achieves up to a 40\% relative reduction in character error rate (CER) on LibriSpeech and generalizes effectively to Mandarin, a language that has remained particularly difficult for prior methods. Code will be released upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Unsupervised Speech Recognition at the Syllable-Level
Wang, Liming
Ni, Junrui
Chang, Kai-Wei
Bhati, Saurabhchand
Harwath, David
Hasegawa-Johnson, Mark
Glass, James R.
Computation and Language
Artificial Intelligence
Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning from non-parallel data. However, existing approaches based on phones often rely on costly resources such as grapheme-to-phoneme converters (G2Ps) and struggle to generalize to languages with ambiguous phoneme boundaries due to training instability. In this paper, we address both challenges by introducing a syllable-level UASR framework based on masked language modeling, which avoids the need for G2P and the instability of GAN-based methods. Our approach achieves up to a 40\% relative reduction in character error rate (CER) on LibriSpeech and generalizes effectively to Mandarin, a language that has remained particularly difficult for prior methods. Code will be released upon acceptance.
title Towards Unsupervised Speech Recognition at the Syllable-Level
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.03639