Classification of Required Data Types and Methods for Text Error Detection and Standardization

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Authors: Hosseini Beheshti, Moluksadat, Abdi Qavidel, Hadi
Format: Recurso digital
Published: Zenodo 2017
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901248293208064
author Hosseini Beheshti, Moluksadat
Abdi Qavidel, Hadi
author_facet Hosseini Beheshti, Moluksadat
Abdi Qavidel, Hadi
contents <p>One of the fundamental stages of automatic text processing is the detection of spelling errors and the standardization of characters. Without passing through this stage, the storage of textual documents faces numerous challenges, leading to disruptions in their machine retrieval. Consequently, specialists in the fields of natural language processing and computational linguistics are continually striving to process various types of data by providing optimal methods and algorithms to achieve standardized data.</p> <p>In English and some other languages, numerous studies have been conducted in this area, and as a result, Persian has also been researched in this context. These various studies have sometimes remained at the level of research and at other times have been presented in the form of products. The present article classifies the various methods and required data types in these studies and specifically describes the process of each, as well as the general approach to measuring their processing accuracy. Additionally, this article describes the functioning of monolingual Persian systems and addresses how they deal with the challenges posed by the Persian language.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_14002415
institution Zenodo
language
publishDate 2017
publisher Zenodo
record_format zenodo
spellingShingle Classification of Required Data Types and Methods for Text Error Detection and Standardization
Hosseini Beheshti, Moluksadat
Abdi Qavidel, Hadi
Spelling Error Detection
<p>One of the fundamental stages of automatic text processing is the detection of spelling errors and the standardization of characters. Without passing through this stage, the storage of textual documents faces numerous challenges, leading to disruptions in their machine retrieval. Consequently, specialists in the fields of natural language processing and computational linguistics are continually striving to process various types of data by providing optimal methods and algorithms to achieve standardized data.</p> <p>In English and some other languages, numerous studies have been conducted in this area, and as a result, Persian has also been researched in this context. These various studies have sometimes remained at the level of research and at other times have been presented in the form of products. The present article classifies the various methods and required data types in these studies and specifically describes the process of each, as well as the general approach to measuring their processing accuracy. Additionally, this article describes the functioning of monolingual Persian systems and addresses how they deal with the challenges posed by the Persian language.</p>
title Classification of Required Data Types and Methods for Text Error Detection and Standardization
topic Spelling Error Detection
url https://doi.org/10.5281/zenodo.14002415