Internal Value Alignment in Large Language Models through Controlled Value Vector Activation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Haoran, Li, Meng, Wang, Xiting, Xu, Zhihao, Huang, Minlie, Jia, Yantao, Lian, Defu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911058048843776
author Jin, Haoran
Li, Meng
Wang, Xiting
Xu, Zhihao
Huang, Minlie
Jia, Yantao
Lian, Defu
author_facet Jin, Haoran
Li, Meng
Wang, Xiting
Xu, Zhihao
Huang, Minlie
Jia, Yantao
Lian, Defu
contents Aligning Large Language Models (LLMs) with human values has attracted increasing attention since it provides clarity, transparency, and the ability to adapt to evolving scenarios. In this paper, we introduce a Controlled Value Vector Activation (ConVA) method that directly aligns the internal values of LLMs by interpreting how a value is encoded in their latent representations and modifies relevant activations to ensure consistent values in LLMs. To ensure an accurate and unbiased interpretation, we propose a context-controlled value vector identification method. To consistently control values without sacrificing model performance, we introduce a gated value vector activation method for effective and minimum degree of value control. Experiments show that our method achieves the highest control success rate across 10 basic values without hurting LLM performance and fluency, and ensures target values even with opposite and potentially malicious input prompts. Source code and data are available at~ https://github.com/hr-jin/ConVA.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
Jin, Haoran
Li, Meng
Wang, Xiting
Xu, Zhihao
Huang, Minlie
Jia, Yantao
Lian, Defu
Computation and Language
Artificial Intelligence
Machine Learning
Aligning Large Language Models (LLMs) with human values has attracted increasing attention since it provides clarity, transparency, and the ability to adapt to evolving scenarios. In this paper, we introduce a Controlled Value Vector Activation (ConVA) method that directly aligns the internal values of LLMs by interpreting how a value is encoded in their latent representations and modifies relevant activations to ensure consistent values in LLMs. To ensure an accurate and unbiased interpretation, we propose a context-controlled value vector identification method. To consistently control values without sacrificing model performance, we introduce a gated value vector activation method for effective and minimum degree of value control. Experiments show that our method achieves the highest control success rate across 10 basic values without hurting LLM performance and fluency, and ensures target values even with opposite and potentially malicious input prompts. Source code and data are available at~ https://github.com/hr-jin/ConVA.
title Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.11316