Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Min, Zhang, Yuchen, Zhang, Shichang, Yin, Fan, Zhang, Dan, Zou, Difan, Yue, Yisong, Hu, Ziniu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914970324697088
author Cai, Min
Zhang, Yuchen
Zhang, Shichang
Yin, Fan
Zhang, Dan
Zou, Difan
Yue, Yisong
Hu, Ziniu
author_facet Cai, Min
Zhang, Yuchen
Zhang, Shichang
Yin, Fan
Zhang, Dan
Zou, Difan
Yue, Yisong
Hu, Ziniu
contents We propose SelfControl, an inference-time model control method utilizing gradients to control the behavior of large language models (LLMs) without explicit human annotations. Given a desired behavior expressed in a natural language suffix string concatenated to the input prompt, SelfControl computes gradients of the LLM's self-evaluation of the suffix with respect to its latent representations. The gradients are used to directly control the auto-regressive generation process towards desired behaviors, which eliminates human supervision, achieves precise and transparent control, and offers on-the-fly adaptability. To further enhance efficiency, we introduce SelfControl_{Prefix}, a compact module that encapsulates the learned representations from gradients into a SelfControl_{Prefix}, facilitating efficient inference-time control with no latency compared to the original model and allowing control for multiple behaviors simultaneously. Our experiments demonstrate SelfControl's efficacy across multiple domains, where it improves over SOTA for 8.3% in detoxification, 3.1% in truthfulness enhancement, 4%~10% in controlling on emotion tones, and 48.2% in privacy protection, i.e., completely remove privacy leakage issue. Additionally, we demonstrate that SelfControl can be used for data synthesis and to improve reasoning abilities.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02721
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
Cai, Min
Zhang, Yuchen
Zhang, Shichang
Yin, Fan
Zhang, Dan
Zou, Difan
Yue, Yisong
Hu, Ziniu
Computation and Language
Artificial Intelligence
We propose SelfControl, an inference-time model control method utilizing gradients to control the behavior of large language models (LLMs) without explicit human annotations. Given a desired behavior expressed in a natural language suffix string concatenated to the input prompt, SelfControl computes gradients of the LLM's self-evaluation of the suffix with respect to its latent representations. The gradients are used to directly control the auto-regressive generation process towards desired behaviors, which eliminates human supervision, achieves precise and transparent control, and offers on-the-fly adaptability. To further enhance efficiency, we introduce SelfControl_{Prefix}, a compact module that encapsulates the learned representations from gradients into a SelfControl_{Prefix}, facilitating efficient inference-time control with no latency compared to the original model and allowing control for multiple behaviors simultaneously. Our experiments demonstrate SelfControl's efficacy across multiple domains, where it improves over SOTA for 8.3% in detoxification, 3.1% in truthfulness enhancement, 4%~10% in controlling on emotion tones, and 48.2% in privacy protection, i.e., completely remove privacy leakage issue. Additionally, we demonstrate that SelfControl can be used for data synthesis and to improve reasoning abilities.
title Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.02721