ReCoVeR the Target Language: Language Steering without Sacrificing Task Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sterz, Hannah, Schmidt, Fabian David, Glavaš, Goran, Vulić, Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909795314827264
author Sterz, Hannah
Schmidt, Fabian David
Glavaš, Goran
Vulić, Ivan
author_facet Sterz, Hannah
Schmidt, Fabian David
Glavaš, Goran
Vulić, Ivan
contents As they become increasingly multilingual, Large Language Models (LLMs) exhibit more language confusion, i.e., they tend to generate answers in a language different from the language of the prompt or the answer language explicitly requested by the user. In this work, we propose ReCoVeR (REducing language COnfusion in VEctor Representations), a novel lightweight approach for reducing language confusion based on language-specific steering vectors. We first isolate language vectors with the help of multi-parallel corpus and then effectively leverage those vectors for effective LLM steering via fixed (i.e., unsupervised) as well as trainable steering functions. Our extensive evaluation, encompassing three benchmarks and 18 languages, shows that ReCoVeR effectively mitigates language confusion in both monolingual and cross-lingual setups while at the same time -- and in contrast to prior language steering methods -- retaining task performance. Our data code is available at https://github.com/hSterz/recover.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14814
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReCoVeR the Target Language: Language Steering without Sacrificing Task Performance
Sterz, Hannah
Schmidt, Fabian David
Glavaš, Goran
Vulić, Ivan
Computation and Language
As they become increasingly multilingual, Large Language Models (LLMs) exhibit more language confusion, i.e., they tend to generate answers in a language different from the language of the prompt or the answer language explicitly requested by the user. In this work, we propose ReCoVeR (REducing language COnfusion in VEctor Representations), a novel lightweight approach for reducing language confusion based on language-specific steering vectors. We first isolate language vectors with the help of multi-parallel corpus and then effectively leverage those vectors for effective LLM steering via fixed (i.e., unsupervised) as well as trainable steering functions. Our extensive evaluation, encompassing three benchmarks and 18 languages, shows that ReCoVeR effectively mitigates language confusion in both monolingual and cross-lingual setups while at the same time -- and in contrast to prior language steering methods -- retaining task performance. Our data code is available at https://github.com/hSterz/recover.
title ReCoVeR the Target Language: Language Steering without Sacrificing Task Performance
topic Computation and Language
url https://arxiv.org/abs/2509.14814