A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yen, Hao, Ku, Pin-Jui, Siniscalchi, Sabato Marco, Lee, Chin-Hui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914030815281152
author Yen, Hao
Ku, Pin-Jui
Siniscalchi, Sabato Marco
Lee, Chin-Hui
author_facet Yen, Hao
Ku, Pin-Jui
Siniscalchi, Sabato Marco
Lee, Chin-Hui
contents We propose a bottom-up framework for automatic speech recognition (ASR) in syllable-based languages by unifying language-universal articulatory attribute modeling with syllable-level prediction. The system first recognizes sequences or lattices of articulatory attributes that serve as a language-universal, interpretable representation of pronunciation, and then transforms them into syllables through a structured knowledge integration process. We introduce two evaluation metrics, namely Pronunciation Error Rate (PrER) and Syllable Homonym Error Rate (SHER), to evaluate the model's ability to capture pronunciation and handle syllable ambiguities. Experimental results on the AISHELL-1 Mandarin corpus demonstrate that the proposed bottom-up framework achieves competitive performance and exhibits better robustness under low-resource conditions compared to the direct syllable prediction model. Furthermore, we investigate the zero-shot cross-lingual transferability on Japanese and demonstrate significant improvements over character- and phoneme-based baselines by 40% error rate reduction.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
Yen, Hao
Ku, Pin-Jui
Siniscalchi, Sabato Marco
Lee, Chin-Hui
Audio and Speech Processing
We propose a bottom-up framework for automatic speech recognition (ASR) in syllable-based languages by unifying language-universal articulatory attribute modeling with syllable-level prediction. The system first recognizes sequences or lattices of articulatory attributes that serve as a language-universal, interpretable representation of pronunciation, and then transforms them into syllables through a structured knowledge integration process. We introduce two evaluation metrics, namely Pronunciation Error Rate (PrER) and Syllable Homonym Error Rate (SHER), to evaluate the model's ability to capture pronunciation and handle syllable ambiguities. Experimental results on the AISHELL-1 Mandarin corpus demonstrate that the proposed bottom-up framework achieves competitive performance and exhibits better robustness under low-resource conditions compared to the direct syllable prediction model. Furthermore, we investigate the zero-shot cross-lingual transferability on Japanese and demonstrate significant improvements over character- and phoneme-based baselines by 40% error rate reduction.
title A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
topic Audio and Speech Processing
url https://arxiv.org/abs/2509.08173