Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909387860213760 |
|---|---|
| author | Islam, Shayekh Bin Tanvir, Ridwanul Hasan Afnan, Sihat |
| author_facet | Islam, Shayekh Bin Tanvir, Ridwanul Hasan Afnan, Sihat |
| contents | Bangla is the seventh most spoken language by a total number of speakers in the world, and yet the development of an automated grammar checker in this language is an understudied problem. Bangla grammatical error detection is a task of detecting sub-strings of a Bangla text that contain grammatical, punctuation, or spelling errors, which is crucial for developing an automated Bangla typing assistant. Our approach involves breaking down the task as a token classification problem and utilizing state-of-the-art transformer-based models. Finally, we combine the output of these models and apply rule-based post-processing to generate a more reliable and comprehensive result. Our system is evaluated on a dataset consisting of over 25,000 texts from various sources. Our best model achieves a Levenshtein distance score of 1.04. Finally, we provide a detailed analysis of different components of our system. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_08344 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification Islam, Shayekh Bin Tanvir, Ridwanul Hasan Afnan, Sihat Computation and Language Machine Learning Bangla is the seventh most spoken language by a total number of speakers in the world, and yet the development of an automated grammar checker in this language is an understudied problem. Bangla grammatical error detection is a task of detecting sub-strings of a Bangla text that contain grammatical, punctuation, or spelling errors, which is crucial for developing an automated Bangla typing assistant. Our approach involves breaking down the task as a token classification problem and utilizing state-of-the-art transformer-based models. Finally, we combine the output of these models and apply rule-based post-processing to generate a more reliable and comprehensive result. Our system is evaluated on a dataset consisting of over 25,000 texts from various sources. Our best model achieves a Levenshtein distance score of 1.04. Finally, we provide a detailed analysis of different components of our system. |
| title | Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2411.08344 |