Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912403800719360 |
|---|---|
| author | Ho, Luong Le, Khanh Pham, Vinh Nguyen, Bao Tran, Tan Chau, Duc |
| author_facet | Ho, Luong Le, Khanh Pham, Vinh Nguyen, Bao Tran, Tan Chau, Duc |
| contents | Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integration of streaming ITN within streaming ASR remains largely unexplored due to challenges in accuracy, efficiency, and adaptability, particularly in low-resource and limited-context scenarios. In this paper, we introduce a streaming pretrained language model for ITN, leveraging pretrained linguistic representations for improved robustness. To address streaming constraints, we propose Dynamic Context-Aware during training and inference, enabling adaptive chunk size adjustments and the integration of right-context information. Experimental results demonstrate that our method achieves accuracy comparable to non-streaming ITN and surpasses existing streaming ITN models on a Vietnamese dataset, all while maintaining low latency, ensuring seamless integration into ASR systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_24229 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization Ho, Luong Le, Khanh Pham, Vinh Nguyen, Bao Tran, Tan Chau, Duc Computation and Language Sound Audio and Speech Processing Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integration of streaming ITN within streaming ASR remains largely unexplored due to challenges in accuracy, efficiency, and adaptability, particularly in low-resource and limited-context scenarios. In this paper, we introduce a streaming pretrained language model for ITN, leveraging pretrained linguistic representations for improved robustness. To address streaming constraints, we propose Dynamic Context-Aware during training and inference, enabling adaptive chunk size adjustments and the integration of right-context information. Experimental results demonstrate that our method achieves accuracy comparable to non-streaming ITN and surpasses existing streaming ITN models on a Vietnamese dataset, all while maintaining low latency, ensuring seamless integration into ASR systems. |
| title | Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.24229 |