Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Veitsman, Yana, Liu, Yihong, Schütze, Hinrich
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911530358931456
author Veitsman, Yana
Liu, Yihong
Schütze, Hinrich
author_facet Veitsman, Yana
Liu, Yihong
Schütze, Hinrich
contents Better cross-lingual alignment is often assumed to yield better cross-lingual transfer. However, explicit alignment techniques -- despite increasing embedding similarity -- frequently fail to improve token-level downstream performance. In this work, we show that this mismatch arises because alignment and downstream task objectives are largely orthogonal, and because the downstream benefits from alignment vary substantially across languages and task types. We analyze four XLM-R encoder models aligned on different language pairs and fine-tuned for either POS Tagging or Sentence Classification. Using representational analyses, including embedding distances, gradient similarities, and gradient magnitudes for both task and alignment losses, we find that: (1) embedding distances alone are unreliable predictors of improvements (or degradations) in task performance and (2) alignment and task gradients are often close to orthogonal, indicating that optimizing one objective may contribute little to optimizing the other. Taken together, our findings explain why ``better'' alignment often fails to translate into ``better'' cross-lingual transfer. Based on these insights, we provide practical guidelines for combining cross-lingual alignment with task-specific fine-tuning, highlighting the importance of careful loss selection.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18863
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
Veitsman, Yana
Liu, Yihong
Schütze, Hinrich
Computation and Language
Better cross-lingual alignment is often assumed to yield better cross-lingual transfer. However, explicit alignment techniques -- despite increasing embedding similarity -- frequently fail to improve token-level downstream performance. In this work, we show that this mismatch arises because alignment and downstream task objectives are largely orthogonal, and because the downstream benefits from alignment vary substantially across languages and task types. We analyze four XLM-R encoder models aligned on different language pairs and fine-tuned for either POS Tagging or Sentence Classification. Using representational analyses, including embedding distances, gradient similarities, and gradient magnitudes for both task and alignment losses, we find that: (1) embedding distances alone are unreliable predictors of improvements (or degradations) in task performance and (2) alignment and task gradients are often close to orthogonal, indicating that optimizing one objective may contribute little to optimizing the other. Taken together, our findings explain why ``better'' alignment often fails to translate into ``better'' cross-lingual transfer. Based on these insights, we provide practical guidelines for combining cross-lingual alignment with task-specific fine-tuning, highlighting the importance of careful loss selection.
title Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
topic Computation and Language
url https://arxiv.org/abs/2603.18863