Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910064236822528 |
|---|---|
| author | Ng, York Hay Khan, Aditya Lu, Xiang Salloum, Matteo Zhou, Michael Hoang, Phuong H. Doğruöz, A. Seza Lee, En-Shiun Annie |
| author_facet | Ng, York Hay Khan, Aditya Lu, Xiang Salloum, Matteo Zhou, Michael Hoang, Phuong H. Doğruöz, A. Seza Lee, En-Shiun Annie |
| contents | Existing linguistic knowledge bases such as URIEL+ provide valuable geographic, genetic and typological distances for cross-lingual transfer but suffer from two key limitations. First, their one-size-fits-all vector representations are ill-suited to the diverse structures of linguistic data. Second, they lack a principled method for aggregating these signals into a single, comprehensive score. In this paper, we address these gaps by introducing a framework for type-matched language distances. We propose novel, structure-aware representations for each distance type: speaker-weighted distributions for geography, hyperbolic embeddings for genealogy, and a latent variables model for typology. We unify these signals into a robust, task-agnostic composite distance. Across multiple zero-shot transfer benchmarks, we demonstrate that our representations significantly improve transfer performance when the distance type is relevant to the task, while our composite distance yields gains in most tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_19217 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+ Ng, York Hay Khan, Aditya Lu, Xiang Salloum, Matteo Zhou, Michael Hoang, Phuong H. Doğruöz, A. Seza Lee, En-Shiun Annie Computation and Language Existing linguistic knowledge bases such as URIEL+ provide valuable geographic, genetic and typological distances for cross-lingual transfer but suffer from two key limitations. First, their one-size-fits-all vector representations are ill-suited to the diverse structures of linguistic data. Second, they lack a principled method for aggregating these signals into a single, comprehensive score. In this paper, we address these gaps by introducing a framework for type-matched language distances. We propose novel, structure-aware representations for each distance type: speaker-weighted distributions for geography, hyperbolic embeddings for genealogy, and a latent variables model for typology. We unify these signals into a robust, task-agnostic composite distance. Across multiple zero-shot transfer benchmarks, we demonstrate that our representations significantly improve transfer performance when the distance type is relevant to the task, while our composite distance yields gains in most tasks. |
| title | Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+ |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.19217 |