Franken-Adapter: Cross-Lingual Adaptation of LLMs by Embedding Surgery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Fan, Yu, Honglin, Chung, Grace, Cohn, Trevor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915146997170176
author Jiang, Fan
Yu, Honglin
Chung, Grace
Cohn, Trevor
author_facet Jiang, Fan
Yu, Honglin
Chung, Grace
Cohn, Trevor
contents The capabilities of Large Language Models (LLMs) in low-resource languages lag far behind those in English, making their universal accessibility a significant challenge. To alleviate this, we present $\textit{Franken-Adapter}$, a modular language adaptation approach for decoder-only LLMs with embedding surgery. Our method begins by creating customized vocabularies for target languages and performing language adaptation through embedding tuning on multilingual data. These pre-trained embeddings are subsequently integrated with LLMs that have been instruction-tuned on English alignment data to enable zero-shot cross-lingual transfer. Our experiments on $\texttt{Gemma2}$ models with up to 27B parameters demonstrate improvements of up to 20% across 96 languages, spanning both discriminative and generative tasks, with minimal regressions ($<$1%) in English. Further in-depth analysis reveals the critical role of customizing tokenizers in enhancing language adaptation, while boosting inference efficiency. Additionally, we show the versatility of our method by achieving a 14% improvement over a math-optimized LLM across 20 languages, offering a modular solution to transfer reasoning abilities across languages post hoc.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08037
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Franken-Adapter: Cross-Lingual Adaptation of LLMs by Embedding Surgery
Jiang, Fan
Yu, Honglin
Chung, Grace
Cohn, Trevor
Computation and Language
The capabilities of Large Language Models (LLMs) in low-resource languages lag far behind those in English, making their universal accessibility a significant challenge. To alleviate this, we present $\textit{Franken-Adapter}$, a modular language adaptation approach for decoder-only LLMs with embedding surgery. Our method begins by creating customized vocabularies for target languages and performing language adaptation through embedding tuning on multilingual data. These pre-trained embeddings are subsequently integrated with LLMs that have been instruction-tuned on English alignment data to enable zero-shot cross-lingual transfer. Our experiments on $\texttt{Gemma2}$ models with up to 27B parameters demonstrate improvements of up to 20% across 96 languages, spanning both discriminative and generative tasks, with minimal regressions ($<$1%) in English. Further in-depth analysis reveals the critical role of customizing tokenizers in enhancing language adaptation, while boosting inference efficiency. Additionally, we show the versatility of our method by achieving a 14% improvement over a math-optimized LLM across 20 languages, offering a modular solution to transfer reasoning abilities across languages post hoc.
title Franken-Adapter: Cross-Lingual Adaptation of LLMs by Embedding Surgery
topic Computation and Language
url https://arxiv.org/abs/2502.08037