Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shipton, Mason, Ng, York Hay, Khan, Aditya, Hoang, Phuong Hanh, Lu, Xiang, Doğruöz, A. Seza, Lee, En-Shiun Annie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918394742177792
author Shipton, Mason
Ng, York Hay
Khan, Aditya
Hoang, Phuong Hanh
Lu, Xiang
Doğruöz, A. Seza
Lee, En-Shiun Annie
author_facet Shipton, Mason
Ng, York Hay
Khan, Aditya
Hoang, Phuong Hanh
Lu, Xiang
Doğruöz, A. Seza
Lee, En-Shiun Annie
contents The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+
Shipton, Mason
Ng, York Hay
Khan, Aditya
Hoang, Phuong Hanh
Lu, Xiang
Doğruöz, A. Seza
Lee, En-Shiun Annie
Computation and Language
The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups.
title Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+
topic Computation and Language
url https://arxiv.org/abs/2510.27183