Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.15539 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912875764776960 |
|---|---|
| author | Keita, Mamadou K. Bremang, Adwoa Le, Huy Owusu, Dennis Homan, Christopher Zampieri, Marcos |
| author_facet | Keita, Mamadou K. Bremang, Adwoa Le, Huy Owusu, Dennis Homan, Christopher Zampieri, Marcos |
| contents | Grammatical error correction (GEC) aims to improve text quality and readability. Previous work on the task focused primarily on high-resource languages, while low-resource languages lack robust tools. To address this shortcoming, we present a study on GEC for Zarma, a language spoken by over five million people in West Africa. We compare three approaches: rule-based methods, machine translation (MT) models, and large language models (LLMs). We evaluated GEC models using a dataset of more than 250,000 examples, including synthetic and human-annotated data. Our results showed that the MT-based approach using M2M100 outperforms others, with a detection rate of 95.82% and a suggestion accuracy of 78.90% in automatic evaluations (AE) and an average score of 3.0 out of 5.0 in manual evaluation (ME) from native speakers for grammar and logical corrections. The rule-based method was effective for spelling errors but failed on complex context-level errors. LLMs -- Gemma 2b and MT5-small -- showed moderate performance. Our work supports use of MT models to enhance GEC in low-resource settings, and we validated these results with Bambara, another West African language. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_15539 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Grammatical Error Correction for Low-Resource Languages: The Case of Zarma Keita, Mamadou K. Bremang, Adwoa Le, Huy Owusu, Dennis Homan, Christopher Zampieri, Marcos Computation and Language Machine Learning Grammatical error correction (GEC) aims to improve text quality and readability. Previous work on the task focused primarily on high-resource languages, while low-resource languages lack robust tools. To address this shortcoming, we present a study on GEC for Zarma, a language spoken by over five million people in West Africa. We compare three approaches: rule-based methods, machine translation (MT) models, and large language models (LLMs). We evaluated GEC models using a dataset of more than 250,000 examples, including synthetic and human-annotated data. Our results showed that the MT-based approach using M2M100 outperforms others, with a detection rate of 95.82% and a suggestion accuracy of 78.90% in automatic evaluations (AE) and an average score of 3.0 out of 5.0 in manual evaluation (ME) from native speakers for grammar and logical corrections. The rule-based method was effective for spelling errors but failed on complex context-level errors. LLMs -- Gemma 2b and MT5-small -- showed moderate performance. Our work supports use of MT models to enhance GEC in low-resource settings, and we validated these results with Bambara, another West African language. |
| title | Grammatical Error Correction for Low-Resource Languages: The Case of Zarma |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2410.15539 |