TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912375473438720 |
|---|---|
| author | Liu, Yutong Xiao, Feng Zhang, Ziyue Yu, Yongbin Huang, Cheng Gao, Fan Wang, Xiangxiang Ban, Ma-bao Fan, Manping Tsering, Thupten Huang, Cheng Luosang, Gadeng Duojie, Renzeng Tashi, Nyima |
| author_facet | Liu, Yutong Xiao, Feng Zhang, Ziyue Yu, Yongbin Huang, Cheng Gao, Fan Wang, Xiangxiang Ban, Ma-bao Fan, Manping Tsering, Thupten Huang, Cheng Luosang, Gadeng Duojie, Renzeng Tashi, Nyima |
| contents | Multi-level Tibetan spelling correction addresses errors at both the character and syllable levels within a unified model. Existing methods focus mainly on single-level correction and lack effective integration of both levels. Moreover, there are no open-source datasets or augmentation methods tailored for this task in Tibetan. To tackle this, we propose a data augmentation approach using unlabeled text to generate multi-level corruptions, and introduce TiSpell, a semi-masked model capable of correcting both character- and syllable-level errors. Although syllable-level correction is more challenging due to its reliance on global context, our semi-masked strategy simplifies this process. We synthesize nine types of corruptions on clean sentences to create a robust training set. Experiments on both simulated and real-world data demonstrate that TiSpell, trained on our dataset, outperforms baseline models and matches the performance of state-of-the-art approaches, confirming its effectiveness. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_08037 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation Liu, Yutong Xiao, Feng Zhang, Ziyue Yu, Yongbin Huang, Cheng Gao, Fan Wang, Xiangxiang Ban, Ma-bao Fan, Manping Tsering, Thupten Huang, Cheng Luosang, Gadeng Duojie, Renzeng Tashi, Nyima Computation and Language Machine Learning Multi-level Tibetan spelling correction addresses errors at both the character and syllable levels within a unified model. Existing methods focus mainly on single-level correction and lack effective integration of both levels. Moreover, there are no open-source datasets or augmentation methods tailored for this task in Tibetan. To tackle this, we propose a data augmentation approach using unlabeled text to generate multi-level corruptions, and introduce TiSpell, a semi-masked model capable of correcting both character- and syllable-level errors. Although syllable-level correction is more challenging due to its reliance on global context, our semi-masked strategy simplifies this process. We synthesize nine types of corruptions on clean sentences to create a robust training set. Experiments on both simulated and real-world data demonstrate that TiSpell, trained on our dataset, outperforms baseline models and matches the performance of state-of-the-art approaches, confirming its effectiveness. |
| title | TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2505.08037 |