Improving Multilingual Language Models by Aligning Representations through Steering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahmoud, Omar, Semage, Buddhika Laknath, Karimpanal, Thommen George, Rana, Santu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912555371331584
author Mahmoud, Omar
Semage, Buddhika Laknath
Karimpanal, Thommen George
Rana, Santu
author_facet Mahmoud, Omar
Semage, Buddhika Laknath
Karimpanal, Thommen George
Rana, Santu
contents This paper investigates how Large Language Models (LLMs) represent non-English tokens -- a question that remains underexplored despite recent progress. We propose a lightweight intervention method using representation steering, where a learned vector is added to the residual stream at a single model layer to enhance multilingual performance. Through extensive experiments across seven competitive baselines -- including prompt optimization, supervised fine-tuning (SFT), in-context learning, cross-lingual transfer, and translation-based methods-we show that our approach consistently outperforms most alternatives. In particular, it achieves performance on par with production-grade translation systems while requiring far fewer resources. We further explore the complementarity between our method and SFT, demonstrating that steering offers a direct, efficient way to realign internal representations. These findings underscore the potential of activation-level interventions as a powerful tool for improving the multilingual capabilities of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Multilingual Language Models by Aligning Representations through Steering
Mahmoud, Omar
Semage, Buddhika Laknath
Karimpanal, Thommen George
Rana, Santu
Computation and Language
This paper investigates how Large Language Models (LLMs) represent non-English tokens -- a question that remains underexplored despite recent progress. We propose a lightweight intervention method using representation steering, where a learned vector is added to the residual stream at a single model layer to enhance multilingual performance. Through extensive experiments across seven competitive baselines -- including prompt optimization, supervised fine-tuning (SFT), in-context learning, cross-lingual transfer, and translation-based methods-we show that our approach consistently outperforms most alternatives. In particular, it achieves performance on par with production-grade translation systems while requiring far fewer resources. We further explore the complementarity between our method and SFT, demonstrating that steering offers a direct, efficient way to realign internal representations. These findings underscore the potential of activation-level interventions as a powerful tool for improving the multilingual capabilities of LLMs.
title Improving Multilingual Language Models by Aligning Representations through Steering
topic Computation and Language
url https://arxiv.org/abs/2505.12584