Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Danni, Niehues, Jan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909631079514112
author Liu, Danni
Niehues, Jan
author_facet Liu, Danni
Niehues, Jan
contents While large language models demonstrate remarkable capabilities at task-specific applications through fine-tuning, extending these benefits across diverse languages is essential for broad accessibility. However, effective cross-lingual transfer is hindered by LLM performance gaps across languages and the scarcity of fine-tuning data in many languages. Through analysis of LLM internal representations from over 1,000+ language pairs, we discover that middle layers exhibit the strongest potential for cross-lingual alignment. Building on this finding, we propose a middle-layer alignment objective integrated into task-specific training. Our experiments on slot filling, machine translation, and structured text generation show consistent improvements in cross-lingual transfer, especially to lower-resource languages. The method is robust to the choice of alignment languages and generalizes to languages unseen during alignment. Furthermore, we show that separately trained alignment modules can be merged with existing task-specific modules, improving cross-lingual capabilities without full re-training. Our code is publicly available (https://github.com/dannigt/mid-align).
format Preprint
id arxiv_https___arxiv_org_abs_2502_14830
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
Liu, Danni
Niehues, Jan
Computation and Language
Artificial Intelligence
While large language models demonstrate remarkable capabilities at task-specific applications through fine-tuning, extending these benefits across diverse languages is essential for broad accessibility. However, effective cross-lingual transfer is hindered by LLM performance gaps across languages and the scarcity of fine-tuning data in many languages. Through analysis of LLM internal representations from over 1,000+ language pairs, we discover that middle layers exhibit the strongest potential for cross-lingual alignment. Building on this finding, we propose a middle-layer alignment objective integrated into task-specific training. Our experiments on slot filling, machine translation, and structured text generation show consistent improvements in cross-lingual transfer, especially to lower-resource languages. The method is robust to the choice of alignment languages and generalizes to languages unseen during alignment. Furthermore, we show that separately trained alignment modules can be merged with existing task-specific modules, improving cross-lingual capabilities without full re-training. Our code is publicly available (https://github.com/dannigt/mid-align).
title Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.14830