Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nie, Shangrui, Mai, Florian, Kaczér, David, Welch, Charles, Zhao, Zhixue, Flek, Lucie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909737863348224
author Nie, Shangrui
Mai, Florian
Kaczér, David
Welch, Charles
Zhao, Zhixue
Flek, Lucie
author_facet Nie, Shangrui
Mai, Florian
Kaczér, David
Welch, Charles
Zhao, Zhixue
Flek, Lucie
contents Large language models implicitly encode preferences over human values, yet steering them often requires large training data. In this work, we investigate a simple approach: Can we reliably modify a model's value system in downstream behavior by training it to answer value survey questions accordingly? We first construct value profiles of several open-source LLMs by asking them to rate a series of value-related descriptions spanning 20 distinct human values, which we use as a baseline for subsequent experiments. We then investigate whether the value system of a model can be governed by fine-tuning on the value surveys. We evaluate the effect of finetuning on the model's behavior in two ways; first, we assess how answers change on in-domain, held-out survey questions. Second, we evaluate whether the model's behavior changes in out-of-domain settings (situational scenarios). To this end, we construct a contextualized moral judgment dataset based on Reddit posts and evaluate changes in the model's behavior in text-based adventure games. We demonstrate that our simple approach can not only change the model's answers to in-domain survey questions, but also produces substantial shifts (value alignment) in implicit downstream task behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11414
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
Nie, Shangrui
Mai, Florian
Kaczér, David
Welch, Charles
Zhao, Zhixue
Flek, Lucie
Computation and Language
Large language models implicitly encode preferences over human values, yet steering them often requires large training data. In this work, we investigate a simple approach: Can we reliably modify a model's value system in downstream behavior by training it to answer value survey questions accordingly? We first construct value profiles of several open-source LLMs by asking them to rate a series of value-related descriptions spanning 20 distinct human values, which we use as a baseline for subsequent experiments. We then investigate whether the value system of a model can be governed by fine-tuning on the value surveys. We evaluate the effect of finetuning on the model's behavior in two ways; first, we assess how answers change on in-domain, held-out survey questions. Second, we evaluate whether the model's behavior changes in out-of-domain settings (situational scenarios). To this end, we construct a contextualized moral judgment dataset based on Reddit posts and evaluate changes in the model's behavior in text-based adventure games. We demonstrate that our simple approach can not only change the model's answers to in-domain survey questions, but also produces substantial shifts (value alignment) in implicit downstream task behavior.
title Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
topic Computation and Language
url https://arxiv.org/abs/2508.11414