Steering Language Models With Activation Engineering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Turner, Alexander Matt, Thiergart, Lisa, Leech, Gavin, Udell, David, Vazquez, Juan J., Mini, Ulisse, MacDiarmid, Monte
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914968657461248
author Turner, Alexander Matt
Thiergart, Lisa
Leech, Gavin
Udell, David
Vazquez, Juan J.
Mini, Ulisse
MacDiarmid, Monte
author_facet Turner, Alexander Matt
Thiergart, Lisa
Leech, Gavin
Udell, David
Vazquez, Juan J.
Mini, Ulisse
MacDiarmid, Monte
contents Prompt engineering and finetuning aim to maximize language model performance on a given metric (like toxicity reduction). However, these methods do not fully elicit a model's capabilities. To reduce this gap, we introduce activation engineering: the inference-time modification of activations in order to control (or steer) model outputs. Specifically, we introduce the Activation Addition (ActAdd) technique, which contrasts the intermediate activations on prompt pairs (such as "Love" versus "Hate") to compute a steering vector (Subramani et al. 2022). By tactically adding in e.g. the "Love" - "Hate" steering vector during the forward pass, we achieve SOTA on negative-to-positive sentiment shift and detoxification using models including LLaMA-3 and OPT. ActAdd yields inference-time control over high-level output properties (like topic and sentiment) while preserving performance on off-target tasks. ActAdd is lightweight: it does not require any machine optimization and works with a single pair of data points, which enables rapid iteration over steering. ActAdd demonstrates the power of activation engineering.
format Preprint
id arxiv_https___arxiv_org_abs_2308_10248
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Steering Language Models With Activation Engineering
Turner, Alexander Matt
Thiergart, Lisa
Leech, Gavin
Udell, David
Vazquez, Juan J.
Mini, Ulisse
MacDiarmid, Monte
Computation and Language
Machine Learning
Prompt engineering and finetuning aim to maximize language model performance on a given metric (like toxicity reduction). However, these methods do not fully elicit a model's capabilities. To reduce this gap, we introduce activation engineering: the inference-time modification of activations in order to control (or steer) model outputs. Specifically, we introduce the Activation Addition (ActAdd) technique, which contrasts the intermediate activations on prompt pairs (such as "Love" versus "Hate") to compute a steering vector (Subramani et al. 2022). By tactically adding in e.g. the "Love" - "Hate" steering vector during the forward pass, we achieve SOTA on negative-to-positive sentiment shift and detoxification using models including LLaMA-3 and OPT. ActAdd yields inference-time control over high-level output properties (like topic and sentiment) while preserving performance on off-target tasks. ActAdd is lightweight: it does not require any machine optimization and works with a single pair of data points, which enables rapid iteration over steering. ActAdd demonstrates the power of activation engineering.
title Steering Language Models With Activation Engineering
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2308.10248