Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tjuka, Annika, Forkel, Robert, Rzymski, Christoph, List, Johann-Mattis
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913999837200384
author Tjuka, Annika
Forkel, Robert
Rzymski, Christoph
List, Johann-Mattis
author_facet Tjuka, Annika
Forkel, Robert
Rzymski, Christoph
List, Johann-Mattis
contents Lexical resources are crucial for cross-linguistic analysis and can provide new insights into computational models for natural language learning. Here, we present an advanced database for comparative studies of words with multiple meanings, a phenomenon known as colexification. The new version includes improvements in the handling, selection and presentation of the data. We compare the new database with previous versions and find that our improvements provide a more balanced sample covering more language families worldwide, with enhanced data quality, given that all word forms are provided in phonetic transcription. We conclude that the new Database of Cross-Linguistic Colexifications has the potential to inspire exciting new studies that link cross-linguistic data to open questions in linguistic typology, historical linguistics, psycholinguistics, and computational linguistics.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11377
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data
Tjuka, Annika
Forkel, Robert
Rzymski, Christoph
List, Johann-Mattis
Computation and Language
Databases
Lexical resources are crucial for cross-linguistic analysis and can provide new insights into computational models for natural language learning. Here, we present an advanced database for comparative studies of words with multiple meanings, a phenomenon known as colexification. The new version includes improvements in the handling, selection and presentation of the data. We compare the new database with previous versions and find that our improvements provide a more balanced sample covering more language families worldwide, with enhanced data quality, given that all word forms are provided in phonetic transcription. We conclude that the new Database of Cross-Linguistic Colexifications has the potential to inspire exciting new studies that link cross-linguistic data to open questions in linguistic typology, historical linguistics, psycholinguistics, and computational linguistics.
title Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data
topic Computation and Language
Databases
url https://arxiv.org/abs/2503.11377