Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dumas, Clément, Wendler, Chris, Veselovsky, Veniamin, Monea, Giovanni, West, Robert
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912448175407104
author Dumas, Clément
Wendler, Chris
Veselovsky, Veniamin
Monea, Giovanni
West, Robert
author_facet Dumas, Clément
Wendler, Chris
Veselovsky, Veniamin
Monea, Giovanni
West, Robert
contents A central question in multilingual language modeling is whether large language models (LLMs) develop a universal concept representation, disentangled from specific languages. In this paper, we address this question by analyzing latent representations (latents) during a word-translation task in transformer-based LLMs. We strategically extract latents from a source translation prompt and insert them into the forward pass on a target translation prompt. By doing so, we find that the output language is encoded in the latent at an earlier layer than the concept to be translated. Building on this insight, we conduct two key experiments. First, we demonstrate that we can change the concept without changing the language and vice versa through activation patching alone. Second, we show that patching with the mean representation of a concept across different languages does not affect the models' ability to translate it, but instead improves it. Finally, we generalize to multi-token generation and demonstrate that the model can generate natural language description of those mean representations. Our results provide evidence for the existence of language-agnostic concept representations within the investigated models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08745
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
Dumas, Clément
Wendler, Chris
Veselovsky, Veniamin
Monea, Giovanni
West, Robert
Computation and Language
Artificial Intelligence
A central question in multilingual language modeling is whether large language models (LLMs) develop a universal concept representation, disentangled from specific languages. In this paper, we address this question by analyzing latent representations (latents) during a word-translation task in transformer-based LLMs. We strategically extract latents from a source translation prompt and insert them into the forward pass on a target translation prompt. By doing so, we find that the output language is encoded in the latent at an earlier layer than the concept to be translated. Building on this insight, we conduct two key experiments. First, we demonstrate that we can change the concept without changing the language and vice versa through activation patching alone. Second, we show that patching with the mean representation of a concept across different languages does not affect the models' ability to translate it, but instead improves it. Finally, we generalize to multi-token generation and demonstrate that the model can generate natural language description of those mean representations. Our results provide evidence for the existence of language-agnostic concept representations within the investigated models.
title Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.08745