CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gurgurov, Daniil, Ghussin, Yusser Al, Baeumel, Tanja, Chou, Cheng-Ting, Schramowski, Patrick, Mosbach, Marius, van Genabith, Josef, Ostermann, Simon
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918285960806400
author Gurgurov, Daniil
Ghussin, Yusser Al
Baeumel, Tanja
Chou, Cheng-Ting
Schramowski, Patrick
Mosbach, Marius
van Genabith, Josef
Ostermann, Simon
author_facet Gurgurov, Daniil
Ghussin, Yusser Al
Baeumel, Tanja
Chou, Cheng-Ting
Schramowski, Patrick
Mosbach, Marius
van Genabith, Josef
Ostermann, Simon
contents Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipulating internal representations during inference, has emerged as a more efficient and interpretable technique for adapting models to a target language. Yet, no dedicated benchmarks or evaluation protocols exist to quantify the effectiveness of steering techniques. We introduce CLaS-Bench, a lightweight parallel-question benchmark for evaluating language-forcing behavior in LLMs across 32 languages, enabling systematic evaluation of multilingual steering methods. We evaluate a broad array of steering techniques, including residual-stream DiffMean interventions, probe-derived directions, language-specific neurons, PCA/LDA vectors, Sparse Autoencoders, and prompting baselines. Steering performance is measured along two axes: language control and semantic relevance, combined into a single harmonic-mean steering score. We find that across languages simple residual-based DiffMean method consistently outperforms all other methods. Moreover, a layer-wise analysis reveals that language-specific structure emerges predominantly in later layers and steering directions cluster based on language family. CLaS-Bench is the first standardized benchmark for multilingual steering, enabling both rigorous scientific analysis of language representations and practical evaluation of steering as a low-cost adaptation alternative.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08331
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
Gurgurov, Daniil
Ghussin, Yusser Al
Baeumel, Tanja
Chou, Cheng-Ting
Schramowski, Patrick
Mosbach, Marius
van Genabith, Josef
Ostermann, Simon
Computation and Language
Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipulating internal representations during inference, has emerged as a more efficient and interpretable technique for adapting models to a target language. Yet, no dedicated benchmarks or evaluation protocols exist to quantify the effectiveness of steering techniques. We introduce CLaS-Bench, a lightweight parallel-question benchmark for evaluating language-forcing behavior in LLMs across 32 languages, enabling systematic evaluation of multilingual steering methods. We evaluate a broad array of steering techniques, including residual-stream DiffMean interventions, probe-derived directions, language-specific neurons, PCA/LDA vectors, Sparse Autoencoders, and prompting baselines. Steering performance is measured along two axes: language control and semantic relevance, combined into a single harmonic-mean steering score. We find that across languages simple residual-based DiffMean method consistently outperforms all other methods. Moreover, a layer-wise analysis reveals that language-specific structure emerges predominantly in later layers and steering directions cluster based on language family. CLaS-Bench is the first standardized benchmark for multilingual steering, enabling both rigorous scientific analysis of language representations and practical evaluation of steering as a low-cost adaptation alternative.
title CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
topic Computation and Language
url https://arxiv.org/abs/2601.08331