Disentangling meaning from language in LLM-based machine translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lasnier, Théo, Zebaze, Armel, Seddah, Djamé, Bawden, Rachel, Sagot, Benoît
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910011648638976
author Lasnier, Théo
Zebaze, Armel
Seddah, Djamé
Bawden, Rachel
Sagot, Benoît
author_facet Lasnier, Théo
Zebaze, Armel
Seddah, Djamé
Bawden, Rachel
Sagot, Benoît
contents Mechanistic Interpretability (MI) seeks to explain how neural networks implement their capabilities, but the scale of Large Language Models (LLMs) has limited prior MI work in Machine Translation (MT) to word-level analyses. We study sentence-level MT from a mechanistic perspective by analyzing attention heads to understand how LLMs internally encode and distribute translation functions. We decompose MT into two subtasks: producing text in the target language (i.e. target language identification) and preserving the input sentence's meaning (i.e. sentence equivalence). Across three families of open-source models and 20 translation directions, we find that distinct, sparse sets of attention heads specialize in each subtask. Based on this insight, we construct subtask-specific steering vectors and show that modifying just 1% of the relevant heads enables instruction-free MT performance comparable to instruction-based prompting, while ablating these heads selectively disrupts their corresponding translation functions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04613
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Disentangling meaning from language in LLM-based machine translation
Lasnier, Théo
Zebaze, Armel
Seddah, Djamé
Bawden, Rachel
Sagot, Benoît
Computation and Language
Mechanistic Interpretability (MI) seeks to explain how neural networks implement their capabilities, but the scale of Large Language Models (LLMs) has limited prior MI work in Machine Translation (MT) to word-level analyses. We study sentence-level MT from a mechanistic perspective by analyzing attention heads to understand how LLMs internally encode and distribute translation functions. We decompose MT into two subtasks: producing text in the target language (i.e. target language identification) and preserving the input sentence's meaning (i.e. sentence equivalence). Across three families of open-source models and 20 translation directions, we find that distinct, sparse sets of attention heads specialize in each subtask. Based on this insight, we construct subtask-specific steering vectors and show that modifying just 1% of the relevant heads enables instruction-free MT performance comparable to instruction-based prompting, while ablating these heads selectively disrupts their corresponding translation functions.
title Disentangling meaning from language in LLM-based machine translation
topic Computation and Language
url https://arxiv.org/abs/2602.04613