Controlling Language and Diffusion Models by Transporting Activations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rodriguez, Pau, Blaas, Arno, Klein, Michal, Zappella, Luca, Apostoloff, Nicholas, Cuturi, Marco, Suau, Xavier
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915030510862336
author Rodriguez, Pau
Blaas, Arno
Klein, Michal
Zappella, Luca
Apostoloff, Nicholas
Cuturi, Marco
Suau, Xavier
author_facet Rodriguez, Pau
Blaas, Arno
Klein, Michal
Zappella, Luca
Apostoloff, Nicholas
Cuturi, Marco
Suau, Xavier
contents The increasing capabilities of large generative models and their ever more widespread deployment have raised concerns about their reliability, safety, and potential misuse. To address these issues, recent works have proposed to control model generation by steering model activations in order to effectively induce or prevent the emergence of concepts or behaviors in the generated output. In this paper we introduce Activation Transport (AcT), a general framework to steer activations guided by optimal transport theory that generalizes many previous activation-steering works. AcT is modality-agnostic and provides fine-grained control over the model behavior with negligible computational overhead, while minimally impacting model abilities. We experimentally show the effectiveness and versatility of our approach by addressing key challenges in large language models (LLMs) and text-to-image diffusion models (T2Is). For LLMs, we show that AcT can effectively mitigate toxicity, induce arbitrary concepts, and increase their truthfulness. In T2Is, we show how AcT enables fine-grained style control and concept negation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23054
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Controlling Language and Diffusion Models by Transporting Activations
Rodriguez, Pau
Blaas, Arno
Klein, Michal
Zappella, Luca
Apostoloff, Nicholas
Cuturi, Marco
Suau, Xavier
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
68T07, 49Q22
I.2.6; I.2.7; I.4.8
The increasing capabilities of large generative models and their ever more widespread deployment have raised concerns about their reliability, safety, and potential misuse. To address these issues, recent works have proposed to control model generation by steering model activations in order to effectively induce or prevent the emergence of concepts or behaviors in the generated output. In this paper we introduce Activation Transport (AcT), a general framework to steer activations guided by optimal transport theory that generalizes many previous activation-steering works. AcT is modality-agnostic and provides fine-grained control over the model behavior with negligible computational overhead, while minimally impacting model abilities. We experimentally show the effectiveness and versatility of our approach by addressing key challenges in large language models (LLMs) and text-to-image diffusion models (T2Is). For LLMs, we show that AcT can effectively mitigate toxicity, induce arbitrary concepts, and increase their truthfulness. In T2Is, we show how AcT enables fine-grained style control and concept negation.
title Controlling Language and Diffusion Models by Transporting Activations
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
68T07, 49Q22
I.2.6; I.2.7; I.4.8
url https://arxiv.org/abs/2410.23054