Towards Unsupervised Speech Recognition at the Syllable-Level
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911191776886784 |
|---|---|
| author | Wang, Liming Ni, Junrui Chang, Kai-Wei Bhati, Saurabhchand Harwath, David Hasegawa-Johnson, Mark Glass, James R. |
| author_facet | Wang, Liming Ni, Junrui Chang, Kai-Wei Bhati, Saurabhchand Harwath, David Hasegawa-Johnson, Mark Glass, James R. |
| contents | Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning from non-parallel data. However, existing approaches based on phones often rely on costly resources such as grapheme-to-phoneme converters (G2Ps) and struggle to generalize to languages with ambiguous phoneme boundaries due to training instability. In this paper, we address both challenges by introducing a syllable-level UASR framework based on masked language modeling, which avoids the need for G2P and the instability of GAN-based methods. Our approach achieves up to a 40\% relative reduction in character error rate (CER) on LibriSpeech and generalizes effectively to Mandarin, a language that has remained particularly difficult for prior methods. Code will be released upon acceptance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_03639 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Unsupervised Speech Recognition at the Syllable-Level Wang, Liming Ni, Junrui Chang, Kai-Wei Bhati, Saurabhchand Harwath, David Hasegawa-Johnson, Mark Glass, James R. Computation and Language Artificial Intelligence Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning from non-parallel data. However, existing approaches based on phones often rely on costly resources such as grapheme-to-phoneme converters (G2Ps) and struggle to generalize to languages with ambiguous phoneme boundaries due to training instability. In this paper, we address both challenges by introducing a syllable-level UASR framework based on masked language modeling, which avoids the need for G2P and the instability of GAN-based methods. Our approach achieves up to a 40\% relative reduction in character error rate (CER) on LibriSpeech and generalizes effectively to Mandarin, a language that has remained particularly difficult for prior methods. Code will be released upon acceptance. |
| title | Towards Unsupervised Speech Recognition at the Syllable-Level |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.03639 |