Leveraging Robust Optimization for LLM Alignment under Distribution Shifts

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Mingye, Liu, Yi, Fu, Zheren, Zhang, Yongdong, Mao, Zhendong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914101855256576
author Zhu, Mingye
Liu, Yi
Fu, Zheren
Zhang, Yongdong
Mao, Zhendong
author_facet Zhu, Mingye
Liu, Yi
Fu, Zheren
Zhang, Yongdong
Mao, Zhendong
contents Preference alignment methods are increasingly critical for steering large language models (LLMs) to generate outputs consistent with human values. While recent approaches often rely on synthetic data generated by LLMs for scalability and cost-efficiency reasons, this reliance can introduce distribution shifts that undermine the nuanced representation of human preferences needed for desirable outputs. In this paper, we propose a novel distribution-aware optimization framework that improves preference alignment despite such shifts. Our approach first leverages well-learned classifiers to assign a calibration value to each training sample, quantifying its alignment with the target human-preferred distribution. These values are then incorporated into a robust optimization objective that minimizes the worst-case loss over regions of the data space most relevant to human preferences. By explicitly focusing optimization on the target distribution, our approach mitigates the impact of distributional mismatch and improves the generation of responses that better reflect intended values.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
Zhu, Mingye
Liu, Yi
Fu, Zheren
Zhang, Yongdong
Mao, Zhendong
Computation and Language
Preference alignment methods are increasingly critical for steering large language models (LLMs) to generate outputs consistent with human values. While recent approaches often rely on synthetic data generated by LLMs for scalability and cost-efficiency reasons, this reliance can introduce distribution shifts that undermine the nuanced representation of human preferences needed for desirable outputs. In this paper, we propose a novel distribution-aware optimization framework that improves preference alignment despite such shifts. Our approach first leverages well-learned classifiers to assign a calibration value to each training sample, quantifying its alignment with the target human-preferred distribution. These values are then incorporated into a robust optimization objective that minimizes the worst-case loss over regions of the data space most relevant to human preferences. By explicitly focusing optimization on the target distribution, our approach mitigates the impact of distributional mismatch and improves the generation of responses that better reflect intended values.
title Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
topic Computation and Language
url https://arxiv.org/abs/2504.05831