MoVoC: Morphology-Aware Subword Construction for Geez Script Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Teklehaymanot, Hailay Kidu, Fazlija, Dren, Nejdl, Wolfgang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915488803586048
author Teklehaymanot, Hailay Kidu
Fazlija, Dren
Nejdl, Wolfgang
author_facet Teklehaymanot, Hailay Kidu
Fazlija, Dren
Nejdl, Wolfgang
contents Subword-based tokenization methods often fail to preserve morphological boundaries, a limitation especially pronounced in low-resource, morphologically complex languages such as those written in the Geez script. To address this, we present MoVoC (Morpheme-aware Subword Vocabulary Construction) and train MoVoC-Tok, a tokenizer that integrates supervised morphological analysis into the subword vocabulary. This hybrid segmentation approach combines morpheme-based and Byte Pair Encoding (BPE) tokens to preserve morphological integrity while maintaining lexical meaning. To tackle resource scarcity, we curate and release manually annotated morpheme data for four Geez script languages and a morpheme-aware vocabulary for two of them. While the proposed tokenization method does not lead to significant gains in automatic translation quality, we observe consistent improvements in intrinsic metrics, MorphoScore, and Boundary Precision, highlighting the value of morphology-aware segmentation in enhancing linguistic fidelity and token efficiency. Our morpheme-annotated datasets and tokenizer will be publicly available to support further research in low-resource, morphologically rich languages. Our code and data are available on GitHub: https://github.com/hailaykidu/MoVoC
format Preprint
id arxiv_https___arxiv_org_abs_2509_08812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
Teklehaymanot, Hailay Kidu
Fazlija, Dren
Nejdl, Wolfgang
Computation and Language
Artificial Intelligence
I.2.7; I.2.6; H.3.3
Subword-based tokenization methods often fail to preserve morphological boundaries, a limitation especially pronounced in low-resource, morphologically complex languages such as those written in the Geez script. To address this, we present MoVoC (Morpheme-aware Subword Vocabulary Construction) and train MoVoC-Tok, a tokenizer that integrates supervised morphological analysis into the subword vocabulary. This hybrid segmentation approach combines morpheme-based and Byte Pair Encoding (BPE) tokens to preserve morphological integrity while maintaining lexical meaning. To tackle resource scarcity, we curate and release manually annotated morpheme data for four Geez script languages and a morpheme-aware vocabulary for two of them. While the proposed tokenization method does not lead to significant gains in automatic translation quality, we observe consistent improvements in intrinsic metrics, MorphoScore, and Boundary Precision, highlighting the value of morphology-aware segmentation in enhancing linguistic fidelity and token efficiency. Our morpheme-annotated datasets and tokenizer will be publicly available to support further research in low-resource, morphologically rich languages. Our code and data are available on GitHub: https://github.com/hailaykidu/MoVoC
title MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
topic Computation and Language
Artificial Intelligence
I.2.7; I.2.6; H.3.3
url https://arxiv.org/abs/2509.08812