Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Shayekh Bin, Tanvir, Ridwanul Hasan, Afnan, Sihat
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909387860213760
author Islam, Shayekh Bin
Tanvir, Ridwanul Hasan
Afnan, Sihat
author_facet Islam, Shayekh Bin
Tanvir, Ridwanul Hasan
Afnan, Sihat
contents Bangla is the seventh most spoken language by a total number of speakers in the world, and yet the development of an automated grammar checker in this language is an understudied problem. Bangla grammatical error detection is a task of detecting sub-strings of a Bangla text that contain grammatical, punctuation, or spelling errors, which is crucial for developing an automated Bangla typing assistant. Our approach involves breaking down the task as a token classification problem and utilizing state-of-the-art transformer-based models. Finally, we combine the output of these models and apply rule-based post-processing to generate a more reliable and comprehensive result. Our system is evaluated on a dataset consisting of over 25,000 texts from various sources. Our best model achieves a Levenshtein distance score of 1.04. Finally, we provide a detailed analysis of different components of our system.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08344
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification
Islam, Shayekh Bin
Tanvir, Ridwanul Hasan
Afnan, Sihat
Computation and Language
Machine Learning
Bangla is the seventh most spoken language by a total number of speakers in the world, and yet the development of an automated grammar checker in this language is an understudied problem. Bangla grammatical error detection is a task of detecting sub-strings of a Bangla text that contain grammatical, punctuation, or spelling errors, which is crucial for developing an automated Bangla typing assistant. Our approach involves breaking down the task as a token classification problem and utilizing state-of-the-art transformer-based models. Finally, we combine the output of these models and apply rule-based post-processing to generate a more reliable and comprehensive result. Our system is evaluated on a dataset consisting of over 25,000 texts from various sources. Our best model achieves a Levenshtein distance score of 1.04. Finally, we provide a detailed analysis of different components of our system.
title Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.08344