Salvato in:
Dettagli Bibliografici
Autori principali: Mamasaidov, Mukhammadsaid, Aral, Azizullah, Shopulatov, Abror, Inomjonov, Mironshoh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2508.14586
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915452716843008
author Mamasaidov, Mukhammadsaid
Aral, Azizullah
Shopulatov, Abror
Inomjonov, Mironshoh
author_facet Mamasaidov, Mukhammadsaid
Aral, Azizullah
Shopulatov, Abror
Inomjonov, Mironshoh
contents Southern Uzbek (uzs) is a Turkic language variety spoken by around 5 million people in Afghanistan and differs significantly from Northern Uzbek (uzn) in phonology, lexicon, and orthography. Despite the large number of speakers, Southern Uzbek is underrepresented in natural language processing. We present new resources for Southern Uzbek machine translation, including a 997-sentence FLORES+ dev set, 39,994 parallel sentences from dictionary, literary, and web sources, and a fine-tuned NLLB-200 model (lutfiy). We also propose a post-processing method for restoring Arabic-script half-space characters, which improves handling of morphological boundaries. All datasets, models, and tools are released publicly to support future work on Southern Uzbek and other low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek
Mamasaidov, Mukhammadsaid
Aral, Azizullah
Shopulatov, Abror
Inomjonov, Mironshoh
Computation and Language
Southern Uzbek (uzs) is a Turkic language variety spoken by around 5 million people in Afghanistan and differs significantly from Northern Uzbek (uzn) in phonology, lexicon, and orthography. Despite the large number of speakers, Southern Uzbek is underrepresented in natural language processing. We present new resources for Southern Uzbek machine translation, including a 997-sentence FLORES+ dev set, 39,994 parallel sentences from dictionary, literary, and web sources, and a fine-tuned NLLB-200 model (lutfiy). We also propose a post-processing method for restoring Arabic-script half-space characters, which improves handling of morphological boundaries. All datasets, models, and tools are released publicly to support future work on Southern Uzbek and other low-resource languages.
title Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek
topic Computation and Language
url https://arxiv.org/abs/2508.14586