ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Ye Bhone, Aung, Thura, Thu, Ye Kyaw, Oo, Thazin Myint
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912729931972608
author Lin, Ye Bhone
Aung, Thura
Thu, Ye Kyaw
Oo, Thazin Myint
author_facet Lin, Ye Bhone
Aung, Thura
Thu, Ye Kyaw
Oo, Thazin Myint
contents This paper investigates sequence-to-sequence Transformer models for automatic speech recognition (ASR) error correction in low-resource Burmese, focusing on different feature integration strategies including IPA and alignment information. To our knowledge, this is the first study addressing ASR error correction specifically for Burmese. We evaluate five ASR backbones and show that our ASR Error Correction (AEC) approaches consistently improve word- and character-level accuracy over baseline outputs. The proposed AEC model, combining IPA and alignment features, reduced the average WER of ASR models from 51.56 to 39.82 before augmentation (and 51.56 to 43.59 after augmentation) and improving chrF++ scores from 0.5864 to 0.627, demonstrating consistent gains over the baseline ASR outputs without AEC. Our results highlight the robustness of AEC and the importance of feature design for improving ASR outputs in low-resource settings.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features
Lin, Ye Bhone
Aung, Thura
Thu, Ye Kyaw
Oo, Thazin Myint
Computation and Language
Machine Learning
Sound
I.2.7; I.2.6
This paper investigates sequence-to-sequence Transformer models for automatic speech recognition (ASR) error correction in low-resource Burmese, focusing on different feature integration strategies including IPA and alignment information. To our knowledge, this is the first study addressing ASR error correction specifically for Burmese. We evaluate five ASR backbones and show that our ASR Error Correction (AEC) approaches consistently improve word- and character-level accuracy over baseline outputs. The proposed AEC model, combining IPA and alignment features, reduced the average WER of ASR models from 51.56 to 39.82 before augmentation (and 51.56 to 43.59 after augmentation) and improving chrF++ scores from 0.5864 to 0.627, demonstrating consistent gains over the baseline ASR outputs without AEC. Our results highlight the robustness of AEC and the importance of feature design for improving ASR outputs in low-resource settings.
title ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features
topic Computation and Language
Machine Learning
Sound
I.2.7; I.2.6
url https://arxiv.org/abs/2511.21088