Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918394742177792 |
|---|---|
| author | Shipton, Mason Ng, York Hay Khan, Aditya Hoang, Phuong Hanh Lu, Xiang Doğruöz, A. Seza Lee, En-Shiun Annie |
| author_facet | Shipton, Mason Ng, York Hay Khan, Aditya Hoang, Phuong Hanh Lu, Xiang Doğruöz, A. Seza Lee, En-Shiun Annie |
| contents | The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_27183 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+ Shipton, Mason Ng, York Hay Khan, Aditya Hoang, Phuong Hanh Lu, Xiang Doğruöz, A. Seza Lee, En-Shiun Annie Computation and Language The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups. |
| title | Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+ |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.27183 |