Exploring language relations through syntactic distances and geographic proximity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De Gregorio, Juan, Toral, Raúl, Sánchez, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917793420541952
author De Gregorio, Juan
Toral, Raúl
Sánchez, David
author_facet De Gregorio, Juan
Toral, Raúl
Sánchez, David
contents Languages are grouped into families that share common linguistic traits. While this approach has been successful in understanding genetic relations between diverse languages, more analyses are needed to accurately quantify their relatedness, especially in less studied linguistic levels such as syntax. Here, we explore linguistic distances using series of parts of speech (POS) extracted from the Universal Dependencies dataset. Within an information-theoretic framework, we show that employing POS trigrams maximizes the possibility of capturing syntactic variations while being at the same time compatible with the amount of available data. Linguistic connections are then established by assessing pairwise distances based on the POS distributions. Intriguingly, our analysis reveals definite clusters that correspond to well known language families and groups, with exceptions explained by distinct morphological typologies. Furthermore, we obtain a significant correlation between language similarity and geographic distance, which underscores the influence of spatial proximity on language kinships.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18430
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring language relations through syntactic distances and geographic proximity
De Gregorio, Juan
Toral, Raúl
Sánchez, David
Computation and Language
Data Analysis, Statistics and Probability
Physics and Society
Applications
Languages are grouped into families that share common linguistic traits. While this approach has been successful in understanding genetic relations between diverse languages, more analyses are needed to accurately quantify their relatedness, especially in less studied linguistic levels such as syntax. Here, we explore linguistic distances using series of parts of speech (POS) extracted from the Universal Dependencies dataset. Within an information-theoretic framework, we show that employing POS trigrams maximizes the possibility of capturing syntactic variations while being at the same time compatible with the amount of available data. Linguistic connections are then established by assessing pairwise distances based on the POS distributions. Intriguingly, our analysis reveals definite clusters that correspond to well known language families and groups, with exceptions explained by distinct morphological typologies. Furthermore, we obtain a significant correlation between language similarity and geographic distance, which underscores the influence of spatial proximity on language kinships.
title Exploring language relations through syntactic distances and geographic proximity
topic Computation and Language
Data Analysis, Statistics and Probability
Physics and Society
Applications
url https://arxiv.org/abs/2403.18430