Automatic Text Normalization in Uzbek: Problems, Tools, and Solutions

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Sobirova, Nazira
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902245332746240
author Sobirova, Nazira
author_facet Sobirova, Nazira
contents <p>In recent years, research in the field of Natural Language Processing (NLP) has increased the demand for automated text analysis across multiple languages, including Uzbek. The multi-form, morphologically complex, and stylistically diverse nature of texts written in Uzbek poses certain challenges for automatic analysis. The central focus of this article is the automatic normalization of Uzbek texts—that is, the process of text normalization. It is dedicated to studying the linguistic and technological issues that arise during automatic text normalization in the Uzbek language. Complex morphological structures, polyform words, dialectal variants, Cyrillic-Latin script differences, and non-standard expressions complicate this process. The results of this research contribute to the deeper digital processing of the Uzbek language and to improving the quality of systems for machine translation, speech-to-text conversion, and text analysis</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15700278
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Automatic Text Normalization in Uzbek: Problems, Tools, and Solutions
Sobirova, Nazira
Uzbek language
text normalization
natural language processing
artificial intelligence
neural networks
rule-based approach
<p>In recent years, research in the field of Natural Language Processing (NLP) has increased the demand for automated text analysis across multiple languages, including Uzbek. The multi-form, morphologically complex, and stylistically diverse nature of texts written in Uzbek poses certain challenges for automatic analysis. The central focus of this article is the automatic normalization of Uzbek texts—that is, the process of text normalization. It is dedicated to studying the linguistic and technological issues that arise during automatic text normalization in the Uzbek language. Complex morphological structures, polyform words, dialectal variants, Cyrillic-Latin script differences, and non-standard expressions complicate this process. The results of this research contribute to the deeper digital processing of the Uzbek language and to improving the quality of systems for machine translation, speech-to-text conversion, and text analysis</p>
title Automatic Text Normalization in Uzbek: Problems, Tools, and Solutions
topic Uzbek language
text normalization
natural language processing
artificial intelligence
neural networks
rule-based approach
url https://doi.org/10.5281/zenodo.15700278