Value Drifts: Tracing Value Alignment During LLM Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhatia, Mehar, Nayak, Shravan, Kamath, Gaurav, Mosbach, Marius, Stańczak, Karolina, Shwartz, Vered, Reddy, Siva
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917050642857984
author Bhatia, Mehar
Nayak, Shravan
Kamath, Gaurav
Mosbach, Marius
Stańczak, Karolina
Shwartz, Vered
Reddy, Siva
author_facet Bhatia, Mehar
Nayak, Shravan
Kamath, Gaurav
Mosbach, Marius
Stańczak, Karolina
Shwartz, Vered
Reddy, Siva
contents As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. Therefore, studying the alignment of LLMs with human values has become a crucial field of inquiry. Prior work, however, mostly focuses on evaluating the alignment of fully trained models, overlooking the training dynamics by which models learn to express human values. In this work, we investigate how and at which stage value alignment arises during the course of a model's post-training. Our analysis disentangles the effects of post-training algorithms and datasets, measuring both the magnitude and time of value drifts during training. Experimenting with Llama-3 and Qwen-3 models of different sizes and popular supervised fine-tuning (SFT) and preference optimization datasets and algorithms, we find that the SFT phase generally establishes a model's values, and subsequent preference optimization rarely re-aligns these values. Furthermore, using a synthetic preference dataset that enables controlled manipulation of values, we find that different preference optimization algorithms lead to different value alignment outcomes, even when preference data is held constant. Our findings provide actionable insights into how values are learned during post-training and help to inform data curation, as well as the selection of models and algorithms for preference optimization to improve model alignment to human values.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26707
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Value Drifts: Tracing Value Alignment During LLM Post-Training
Bhatia, Mehar
Nayak, Shravan
Kamath, Gaurav
Mosbach, Marius
Stańczak, Karolina
Shwartz, Vered
Reddy, Siva
Computation and Language
Computers and Society
Machine Learning
As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. Therefore, studying the alignment of LLMs with human values has become a crucial field of inquiry. Prior work, however, mostly focuses on evaluating the alignment of fully trained models, overlooking the training dynamics by which models learn to express human values. In this work, we investigate how and at which stage value alignment arises during the course of a model's post-training. Our analysis disentangles the effects of post-training algorithms and datasets, measuring both the magnitude and time of value drifts during training. Experimenting with Llama-3 and Qwen-3 models of different sizes and popular supervised fine-tuning (SFT) and preference optimization datasets and algorithms, we find that the SFT phase generally establishes a model's values, and subsequent preference optimization rarely re-aligns these values. Furthermore, using a synthetic preference dataset that enables controlled manipulation of values, we find that different preference optimization algorithms lead to different value alignment outcomes, even when preference data is held constant. Our findings provide actionable insights into how values are learned during post-training and help to inform data curation, as well as the selection of models and algorithms for preference optimization to improve model alignment to human values.
title Value Drifts: Tracing Value Alignment During LLM Post-Training
topic Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2510.26707