| _version_ | 1866901248293208064 |
|---|---|
| author | Hosseini Beheshti, Moluksadat Abdi Qavidel, Hadi |
| author_facet | Hosseini Beheshti, Moluksadat Abdi Qavidel, Hadi |
| contents | <p>One of the fundamental stages of automatic text processing is the detection of spelling errors and the standardization of characters. Without passing through this stage, the storage of textual documents faces numerous challenges, leading to disruptions in their machine retrieval. Consequently, specialists in the fields of natural language processing and computational linguistics are continually striving to process various types of data by providing optimal methods and algorithms to achieve standardized data.</p> <p>In English and some other languages, numerous studies have been conducted in this area, and as a result, Persian has also been researched in this context. These various studies have sometimes remained at the level of research and at other times have been presented in the form of products. The present article classifies the various methods and required data types in these studies and specifically describes the process of each, as well as the general approach to measuring their processing accuracy. Additionally, this article describes the functioning of monolingual Persian systems and addresses how they deal with the challenges posed by the Persian language.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_14002415 |
| institution | Zenodo |
| language | |
| publishDate | 2017 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Classification of Required Data Types and Methods for Text Error Detection and Standardization Hosseini Beheshti, Moluksadat Abdi Qavidel, Hadi Spelling Error Detection <p>One of the fundamental stages of automatic text processing is the detection of spelling errors and the standardization of characters. Without passing through this stage, the storage of textual documents faces numerous challenges, leading to disruptions in their machine retrieval. Consequently, specialists in the fields of natural language processing and computational linguistics are continually striving to process various types of data by providing optimal methods and algorithms to achieve standardized data.</p> <p>In English and some other languages, numerous studies have been conducted in this area, and as a result, Persian has also been researched in this context. These various studies have sometimes remained at the level of research and at other times have been presented in the form of products. The present article classifies the various methods and required data types in these studies and specifically describes the process of each, as well as the general approach to measuring their processing accuracy. Additionally, this article describes the functioning of monolingual Persian systems and addresses how they deal with the challenges posed by the Persian language.</p> |
| title | Classification of Required Data Types and Methods for Text Error Detection and Standardization |
| topic | Spelling Error Detection |
| url | https://doi.org/10.5281/zenodo.14002415 |