Can Embedding Similarity Predict Cross-Lingual Transfer? A Systematic Study on African Languages
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912805202952192 |
|---|---|
| author | Idris, Tewodros Kederalah Mitra, Prasenjit Eiselen, Roald |
| author_facet | Idris, Tewodros Kederalah Mitra, Prasenjit Eiselen, Roald |
| contents | Cross-lingual transfer is essential for building NLP systems for low-resource African languages, but practitioners lack reliable methods for selecting source languages. We systematically evaluate five embedding similarity metrics across 816 transfer experiments spanning three NLP tasks, three African-centric multilingual models, and 12 languages from four language families. We find that cosine gap and retrieval-based metrics (P@1, CSLS) reliably predict transfer success ($ρ= 0.4-0.6$), while CKA shows negligible predictive power ($ρ\approx 0.1$). Critically, correlation signs reverse when pooling across models (Simpson's Paradox), so practitioners must validate per-model. Embedding metrics achieve comparable predictive power to URIEL linguistic typology. Our results provide concrete guidance for source language selection and highlight the importance of model-specific analysis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_03168 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Can Embedding Similarity Predict Cross-Lingual Transfer? A Systematic Study on African Languages Idris, Tewodros Kederalah Mitra, Prasenjit Eiselen, Roald Computation and Language Machine Learning Cross-lingual transfer is essential for building NLP systems for low-resource African languages, but practitioners lack reliable methods for selecting source languages. We systematically evaluate five embedding similarity metrics across 816 transfer experiments spanning three NLP tasks, three African-centric multilingual models, and 12 languages from four language families. We find that cosine gap and retrieval-based metrics (P@1, CSLS) reliably predict transfer success ($ρ= 0.4-0.6$), while CKA shows negligible predictive power ($ρ\approx 0.1$). Critically, correlation signs reverse when pooling across models (Simpson's Paradox), so practitioners must validate per-model. Embedding metrics achieve comparable predictive power to URIEL linguistic typology. Our results provide concrete guidance for source language selection and highlight the importance of model-specific analysis. |
| title | Can Embedding Similarity Predict Cross-Lingual Transfer? A Systematic Study on African Languages |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2601.03168 |