Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Nghia Hieu, Hoang, Quan Ngoc, Nguyen, Long Hoang Huu, Van Nguyen, Kiet, Nguyen, Ngan Luu-Thuy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913166510784512
author Nguyen, Nghia Hieu
Hoang, Quan Ngoc
Nguyen, Long Hoang Huu
Van Nguyen, Kiet
Nguyen, Ngan Luu-Thuy
author_facet Nguyen, Nghia Hieu
Hoang, Quan Ngoc
Nguyen, Long Hoang Huu
Van Nguyen, Kiet
Nguyen, Ngan Luu-Thuy
contents Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or words. Although effective, such representations do not explicitly reflect the phonetic structure of speech and often require large vocabularies to maintain adequate coverage. In this work, we are motivated from the phonemic features of Vietnamese to propose a Syllabic-Structure Decoder for ASR, which models speech at the phoneme level instead of the orthographic level. Our approach explicitly captures the phonological composition of syllables, enabling the decoder to generate valid syllabic structures from a compact phonemic inventory. This design more closely aligns with the phonetic realization of speech while significantly reducing vocabulary size. Experimental results on two benchmarks: LSVSC, representing standard speech, and UIT-ViMD, a multi-dialect corpus containing diverse regional pronunciations, show that our method consistently outperforms strong previous baselines, especially pretrained baselines such as PhoWhisper and Wav2Vec2, despite using a substantially smaller vocabulary and no additional training resources. These results highlight the effectiveness of phoneme-based syllabic modeling for ASR in this language. Code for experimental reproducibility will be publicly available upon the acceptance of this paper.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27874
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese
Nguyen, Nghia Hieu
Hoang, Quan Ngoc
Nguyen, Long Hoang Huu
Van Nguyen, Kiet
Nguyen, Ngan Luu-Thuy
Computation and Language
Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or words. Although effective, such representations do not explicitly reflect the phonetic structure of speech and often require large vocabularies to maintain adequate coverage. In this work, we are motivated from the phonemic features of Vietnamese to propose a Syllabic-Structure Decoder for ASR, which models speech at the phoneme level instead of the orthographic level. Our approach explicitly captures the phonological composition of syllables, enabling the decoder to generate valid syllabic structures from a compact phonemic inventory. This design more closely aligns with the phonetic realization of speech while significantly reducing vocabulary size. Experimental results on two benchmarks: LSVSC, representing standard speech, and UIT-ViMD, a multi-dialect corpus containing diverse regional pronunciations, show that our method consistently outperforms strong previous baselines, especially pretrained baselines such as PhoWhisper and Wav2Vec2, despite using a substantially smaller vocabulary and no additional training resources. These results highlight the effectiveness of phoneme-based syllabic modeling for ASR in this language. Code for experimental reproducibility will be publicly available upon the acceptance of this paper.
title Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese
topic Computation and Language
url https://arxiv.org/abs/2605.27874