Activation Scaling for Steering and Interpreting Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Stoehr, Niklas, Du, Kevin, Snæbjarnarson, Vésteinn, West, Robert, Cotterell, Ryan, Schein, Aaron
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916425338191872
author Stoehr, Niklas
Du, Kevin
Snæbjarnarson, Vésteinn
West, Robert
Cotterell, Ryan
Schein, Aaron
author_facet Stoehr, Niklas
Du, Kevin
Snæbjarnarson, Vésteinn
West, Robert
Cotterell, Ryan
Schein, Aaron
contents Given the prompt "Rome is in", can we steer a language model to flip its prediction of an incorrect token "France" to a correct token "Italy" by only multiplying a few relevant activation vectors with scalars? We argue that successfully intervening on a model is a prerequisite for interpreting its internal workings. Concretely, we establish a three-term objective: a successful intervention should flip the correct with the wrong token and vice versa (effectiveness), and leave other tokens unaffected (faithfulness), all while being sparse (minimality). Using gradient-based optimization, this objective lets us learn (and later evaluate) a specific kind of efficient and interpretable intervention: activation scaling only modifies the signed magnitude of activation vectors to strengthen, weaken, or reverse the steering directions already encoded in the model. On synthetic tasks, this intervention performs comparably with steering vectors in terms of effectiveness and faithfulness, but is much more minimal allowing us to pinpoint interpretable model components. We evaluate activation scaling from different angles, compare performance on different datasets, and make activation scalars a learnable function of the activation vectors themselves to generalize to varying-length prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04962
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Activation Scaling for Steering and Interpreting Language Models
Stoehr, Niklas
Du, Kevin
Snæbjarnarson, Vésteinn
West, Robert
Cotterell, Ryan
Schein, Aaron
Computation and Language
Artificial Intelligence
Given the prompt "Rome is in", can we steer a language model to flip its prediction of an incorrect token "France" to a correct token "Italy" by only multiplying a few relevant activation vectors with scalars? We argue that successfully intervening on a model is a prerequisite for interpreting its internal workings. Concretely, we establish a three-term objective: a successful intervention should flip the correct with the wrong token and vice versa (effectiveness), and leave other tokens unaffected (faithfulness), all while being sparse (minimality). Using gradient-based optimization, this objective lets us learn (and later evaluate) a specific kind of efficient and interpretable intervention: activation scaling only modifies the signed magnitude of activation vectors to strengthen, weaken, or reverse the steering directions already encoded in the model. On synthetic tasks, this intervention performs comparably with steering vectors in terms of effectiveness and faithfulness, but is much more minimal allowing us to pinpoint interpretable model components. We evaluate activation scaling from different angles, compare performance on different datasets, and make activation scalars a learnable function of the activation vectors themselves to generalize to varying-length prompts.
title Activation Scaling for Steering and Interpreting Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.04962