TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Yutong, Xiao, Feng, Zhang, Ziyue, Yu, Yongbin, Huang, Cheng, Gao, Fan, Wang, Xiangxiang, Ban, Ma-bao, Fan, Manping, Tsering, Thupten, Luosang, Gadeng, Duojie, Renzeng, Tashi, Nyima
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912375473438720
author Liu, Yutong
Xiao, Feng
Zhang, Ziyue
Yu, Yongbin
Huang, Cheng
Gao, Fan
Wang, Xiangxiang
Ban, Ma-bao
Fan, Manping
Tsering, Thupten
Huang, Cheng
Luosang, Gadeng
Duojie, Renzeng
Tashi, Nyima
author_facet Liu, Yutong
Xiao, Feng
Zhang, Ziyue
Yu, Yongbin
Huang, Cheng
Gao, Fan
Wang, Xiangxiang
Ban, Ma-bao
Fan, Manping
Tsering, Thupten
Huang, Cheng
Luosang, Gadeng
Duojie, Renzeng
Tashi, Nyima
contents Multi-level Tibetan spelling correction addresses errors at both the character and syllable levels within a unified model. Existing methods focus mainly on single-level correction and lack effective integration of both levels. Moreover, there are no open-source datasets or augmentation methods tailored for this task in Tibetan. To tackle this, we propose a data augmentation approach using unlabeled text to generate multi-level corruptions, and introduce TiSpell, a semi-masked model capable of correcting both character- and syllable-level errors. Although syllable-level correction is more challenging due to its reliance on global context, our semi-masked strategy simplifies this process. We synthesize nine types of corruptions on clean sentences to create a robust training set. Experiments on both simulated and real-world data demonstrate that TiSpell, trained on our dataset, outperforms baseline models and matches the performance of state-of-the-art approaches, confirming its effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08037
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation
Liu, Yutong
Xiao, Feng
Zhang, Ziyue
Yu, Yongbin
Huang, Cheng
Gao, Fan
Wang, Xiangxiang
Ban, Ma-bao
Fan, Manping
Tsering, Thupten
Huang, Cheng
Luosang, Gadeng
Duojie, Renzeng
Tashi, Nyima
Computation and Language
Machine Learning
Multi-level Tibetan spelling correction addresses errors at both the character and syllable levels within a unified model. Existing methods focus mainly on single-level correction and lack effective integration of both levels. Moreover, there are no open-source datasets or augmentation methods tailored for this task in Tibetan. To tackle this, we propose a data augmentation approach using unlabeled text to generate multi-level corruptions, and introduce TiSpell, a semi-masked model capable of correcting both character- and syllable-level errors. Although syllable-level correction is more challenging due to its reliance on global context, our semi-masked strategy simplifies this process. We synthesize nine types of corruptions on clean sentences to create a robust training set. Experiments on both simulated and real-world data demonstrate that TiSpell, trained on our dataset, outperforms baseline models and matches the performance of state-of-the-art approaches, confirming its effectiveness.
title TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.08037