Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Botong, Li, Shuo, Hounie, Ignacio, Bastani, Osbert, Ding, Dongsheng, Ribeiro, Alejandro
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2505.19387
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911287909285888
author Zhang, Botong
Li, Shuo
Hounie, Ignacio
Bastani, Osbert
Ding, Dongsheng
Ribeiro, Alejandro
author_facet Zhang, Botong
Li, Shuo
Hounie, Ignacio
Bastani, Osbert
Ding, Dongsheng
Ribeiro, Alejandro
contents We study the problem of computing an optimal large language model (LLM) policy for the constrained alignment problem, where the goal is to maximize a primary reward objective while satisfying constraints on secondary utilities. Despite the popularity of Lagrangian-based LLM policy search in constrained alignment, iterative primal-dual methods often fail to converge, and non-iterative dual-based methods do not achieve optimality in the LLM parameter space. To address these challenges, we employ Lagrangian duality to develop an iterative dual-based alignment method that alternates between updating the LLM policy via Lagrangian maximization and updating the dual variable via dual descent. In theory, we characterize the primal-dual gap between the primal value in the distribution space and the dual value in the LLM parameter space. We further quantify the optimality gap of the learned LLM policies at near-optimal dual variables with respect to both the objective and the constraint functions. These results prove that dual-based alignment methods can find an optimal constrained LLM policy, up to an LLM parametrization gap. We demonstrate the effectiveness and merits of our approach through extensive experiments conducted on the PKU-SafeRLHF and Anthropic HH-RLHF datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Alignment of large language models with constrained learning
Zhang, Botong
Li, Shuo
Hounie, Ignacio
Bastani, Osbert
Ding, Dongsheng
Ribeiro, Alejandro
Machine Learning
Systems and Control
Optimization and Control
We study the problem of computing an optimal large language model (LLM) policy for the constrained alignment problem, where the goal is to maximize a primary reward objective while satisfying constraints on secondary utilities. Despite the popularity of Lagrangian-based LLM policy search in constrained alignment, iterative primal-dual methods often fail to converge, and non-iterative dual-based methods do not achieve optimality in the LLM parameter space. To address these challenges, we employ Lagrangian duality to develop an iterative dual-based alignment method that alternates between updating the LLM policy via Lagrangian maximization and updating the dual variable via dual descent. In theory, we characterize the primal-dual gap between the primal value in the distribution space and the dual value in the LLM parameter space. We further quantify the optimality gap of the learned LLM policies at near-optimal dual variables with respect to both the objective and the constraint functions. These results prove that dual-based alignment methods can find an optimal constrained LLM policy, up to an LLM parametrization gap. We demonstrate the effectiveness and merits of our approach through extensive experiments conducted on the PKU-SafeRLHF and Anthropic HH-RLHF datasets.
title Alignment of large language models with constrained learning
topic Machine Learning
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2505.19387