Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dong, Lukuang, Li, Ziwei, Yusuyin, Saierdaer, Zhao, Xianyu, Ou, Zhijian
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908924234432512
author Dong, Lukuang
Li, Ziwei
Yusuyin, Saierdaer
Zhao, Xianyu
Ou, Zhijian
author_facet Dong, Lukuang
Li, Ziwei
Yusuyin, Saierdaer
Zhao, Xianyu
Ou, Zhijian
contents Phoneme-based ASR factorizes recognition into speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G), enabling cross-lingual acoustic sharing while keeping language-specific orthography in a separate module. While large language models (LLMs) are promising for P2G, multilingual P2G remains challenging due to language-aware generation and severe cross-language data imbalance. We study multilingual LLM-based P2G on the ten-language CV-Lang10 benchmark. We examine robustness strategies that account for S2P uncertainty, including DANP and Simplified SKM (S-SKM). S-SKM is a Monte Carlo approximation that avoids CTC-based S2P probability weighting in P2G training. Robust training and low-resource oversampling reduce the average WER from 10.56% to 7.66%.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29217
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
Dong, Lukuang
Li, Ziwei
Yusuyin, Saierdaer
Zhao, Xianyu
Ou, Zhijian
Audio and Speech Processing
Computation and Language
Sound
Phoneme-based ASR factorizes recognition into speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G), enabling cross-lingual acoustic sharing while keeping language-specific orthography in a separate module. While large language models (LLMs) are promising for P2G, multilingual P2G remains challenging due to language-aware generation and severe cross-language data imbalance. We study multilingual LLM-based P2G on the ten-language CV-Lang10 benchmark. We examine robustness strategies that account for S2P uncertainty, including DANP and Simplified SKM (S-SKM). S-SKM is a Monte Carlo approximation that avoids CTC-based S2P probability weighting in P2G training. Robust training and low-resource oversampling reduce the average WER from 10.56% to 7.66%.
title Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2603.29217