Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Rongzhi, Ye, Liqin, Heng, Yuzhao, Chen, Xiang, Yu, Tong, Kong, Lingkai, Chava, Sudheer, Zhang, Chao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914335882739712
author Zhang, Rongzhi
Ye, Liqin
Heng, Yuzhao
Chen, Xiang
Yu, Tong
Kong, Lingkai
Chava, Sudheer
Zhang, Chao
author_facet Zhang, Rongzhi
Ye, Liqin
Heng, Yuzhao
Chen, Xiang
Yu, Tong
Kong, Lingkai
Chava, Sudheer
Zhang, Chao
contents Precise attribute intensity control--generating Large Language Model (LLM) outputs with specific, user-defined attribute intensities--is crucial for AI systems adaptable to diverse user expectations. Current LLM alignment methods, however, typically provide only directional or open-ended guidance, failing to reliably achieve exact attribute intensities. We address this limitation with three key designs: (1) reformulating precise attribute intensity control as a target-reaching problem, rather than simple maximization; (2) training a lightweight value function via temporal-difference learning to predict final attribute intensity scores from partial generations, thereby steering LLM outputs; and (3) employing gradient-based interventions on hidden representations to navigate the model precisely towards specific attribute intensity targets. Our method enables fine-grained, continuous control over attribute intensities, moving beyond simple directional alignment. Experiments on LLaMA-3.2-3b and Phi-4-mini confirm our method's ability to steer text generation to user-specified attribute intensities with high accuracy. Finally, we demonstrate efficiency enhancements across three downstream tasks: preference data synthesis, Pareto frontier approximation and optimization, and distillation of aligned behaviors for intervention-free inference. Our code is available on https://github.com/Pre-Control/pre-control
format Preprint
id arxiv_https___arxiv_org_abs_2510_12121
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
Zhang, Rongzhi
Ye, Liqin
Heng, Yuzhao
Chen, Xiang
Yu, Tong
Kong, Lingkai
Chava, Sudheer
Zhang, Chao
Artificial Intelligence
Computation and Language
Machine Learning
Precise attribute intensity control--generating Large Language Model (LLM) outputs with specific, user-defined attribute intensities--is crucial for AI systems adaptable to diverse user expectations. Current LLM alignment methods, however, typically provide only directional or open-ended guidance, failing to reliably achieve exact attribute intensities. We address this limitation with three key designs: (1) reformulating precise attribute intensity control as a target-reaching problem, rather than simple maximization; (2) training a lightweight value function via temporal-difference learning to predict final attribute intensity scores from partial generations, thereby steering LLM outputs; and (3) employing gradient-based interventions on hidden representations to navigate the model precisely towards specific attribute intensity targets. Our method enables fine-grained, continuous control over attribute intensities, moving beyond simple directional alignment. Experiments on LLaMA-3.2-3b and Phi-4-mini confirm our method's ability to steer text generation to user-specified attribute intensities with high accuracy. Finally, we demonstrate efficiency enhancements across three downstream tasks: preference data synthesis, Pareto frontier approximation and optimization, and distillation of aligned behaviors for intervention-free inference. Our code is available on https://github.com/Pre-Control/pre-control
title Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.12121