VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Woojin, Hyeon, Sieun, Oh, Jusang, Do, Jaeyoung
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910009995034624
author Kim, Woojin
Hyeon, Sieun
Oh, Jusang
Do, Jaeyoung
author_facet Kim, Woojin
Hyeon, Sieun
Oh, Jusang
Do, Jaeyoung
contents Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path, yet three gaps persist: extraction often ignores hierarchical structure, evaluation detects presence but not calibrated intensity, and the steerability of LLMs at controlled intensities remains insufficiently understood. To address these limitations, we introduce VALUEFLOW, the first unified framework that spans extraction, evaluation, and steering with calibrated intensity control. The framework integrates three components: (i) HIVES, a hierarchical value embedding space that captures intra- and cross-theory value structure; (ii) the Value Intensity DataBase (VIDB), a large-scale resource of value-labeled texts with intensity estimates derived from ranking-based aggregation; and (iii) an anchor-based evaluator that produces consistent intensity scores for model outputs by ranking them against VIDB panels. Using VALUEFLOW, we conduct a comprehensive large-scale study across ten models and four value theories, identifying asymmetries in steerability and composition laws for multi-value control. This paper establishes a scalable infrastructure for evaluating and controlling value intensity, advancing pluralistic alignment of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03160
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
Kim, Woojin
Hyeon, Sieun
Oh, Jusang
Do, Jaeyoung
Artificial Intelligence
Computation and Language
Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path, yet three gaps persist: extraction often ignores hierarchical structure, evaluation detects presence but not calibrated intensity, and the steerability of LLMs at controlled intensities remains insufficiently understood. To address these limitations, we introduce VALUEFLOW, the first unified framework that spans extraction, evaluation, and steering with calibrated intensity control. The framework integrates three components: (i) HIVES, a hierarchical value embedding space that captures intra- and cross-theory value structure; (ii) the Value Intensity DataBase (VIDB), a large-scale resource of value-labeled texts with intensity estimates derived from ranking-based aggregation; and (iii) an anchor-based evaluator that produces consistent intensity scores for model outputs by ranking them against VIDB panels. Using VALUEFLOW, we conduct a comprehensive large-scale study across ten models and four value theories, identifying asymmetries in steerability and composition laws for multi-value control. This paper establishes a scalable infrastructure for evaluating and controlling value intensity, advancing pluralistic alignment of LLMs.
title VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.03160