Annotating and Inferring Compositional Structures in Numeral Systems Across Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rubehn, Arne, Rzymski, Christoph, Ciucci, Luca, van Dam, Kellen Parker, Kučerová, Alžběta, Bocklage, Katja, Snee, David, Stephen, Abishek, List, Johann-Mattis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909524759150592
author Rubehn, Arne
Rzymski, Christoph
Ciucci, Luca
van Dam, Kellen Parker
Kučerová, Alžběta
Bocklage, Katja
Snee, David
Stephen, Abishek
List, Johann-Mattis
author_facet Rubehn, Arne
Rzymski, Christoph
Ciucci, Luca
van Dam, Kellen Parker
Kučerová, Alžběta
Bocklage, Katja
Snee, David
Stephen, Abishek
List, Johann-Mattis
contents Numeral systems across the world's languages vary in fascinating ways, both regarding their synchronic structure and the diachronic processes that determined how they evolved in their current shape. For a proper comparison of numeral systems across different languages, however, it is important to code them in a standardized form that allows for the comparison of basic properties. Here, we present a simple but effective coding scheme for numeral annotation, along with a workflow that helps to code numeral systems in a computer-assisted manner, providing sample data for numerals from 1 to 40 in 25 typologically diverse languages. We perform a thorough analysis of the sample, focusing on the systematic comparison between the underlying and the surface morphological structure. We further experiment with automated models for morpheme segmentation, where we find allomorphy as the major reason for segmentation errors. Finally, we show that subword tokenization algorithms are not viable for discovering morphemes in low-resource scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Annotating and Inferring Compositional Structures in Numeral Systems Across Languages
Rubehn, Arne
Rzymski, Christoph
Ciucci, Luca
van Dam, Kellen Parker
Kučerová, Alžběta
Bocklage, Katja
Snee, David
Stephen, Abishek
List, Johann-Mattis
Computation and Language
J.5
Numeral systems across the world's languages vary in fascinating ways, both regarding their synchronic structure and the diachronic processes that determined how they evolved in their current shape. For a proper comparison of numeral systems across different languages, however, it is important to code them in a standardized form that allows for the comparison of basic properties. Here, we present a simple but effective coding scheme for numeral annotation, along with a workflow that helps to code numeral systems in a computer-assisted manner, providing sample data for numerals from 1 to 40 in 25 typologically diverse languages. We perform a thorough analysis of the sample, focusing on the systematic comparison between the underlying and the surface morphological structure. We further experiment with automated models for morpheme segmentation, where we find allomorphy as the major reason for segmentation errors. Finally, we show that subword tokenization algorithms are not viable for discovering morphemes in low-resource scenarios.
title Annotating and Inferring Compositional Structures in Numeral Systems Across Languages
topic Computation and Language
J.5
url https://arxiv.org/abs/2503.01625