One-Shot Safety Alignment for Large Language Models via Optimal Dualization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Xinmeng, Li, Shuo, Dobriban, Edgar, Bastani, Osbert, Hassani, Hamed, Ding, Dongsheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913583020900352
author Huang, Xinmeng
Li, Shuo
Dobriban, Edgar
Bastani, Osbert
Hassani, Hamed
Ding, Dongsheng
author_facet Huang, Xinmeng
Li, Shuo
Dobriban, Edgar
Bastani, Osbert
Hassani, Hamed
Ding, Dongsheng
contents The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safety. A promising approach is to enforce safety constraints through Reinforcement Learning from Human Feedback (RLHF). For such constrained RLHF, typical Lagrangian-based primal-dual policy optimization methods are computationally expensive and often unstable. This paper presents a perspective of dualization that reduces constrained alignment to an equivalent unconstrained alignment problem. We do so by pre-optimizing a smooth and convex dual function that has a closed form. This shortcut eliminates the need for cumbersome primal-dual policy iterations, greatly reducing the computational burden and improving training stability. Our strategy leads to two practical algorithms in model-based and preference-based settings (MoCAN and PeCAN, respectively). A broad range of experiments demonstrate the effectiveness and merits of our algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19544
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle One-Shot Safety Alignment for Large Language Models via Optimal Dualization
Huang, Xinmeng
Li, Shuo
Dobriban, Edgar
Bastani, Osbert
Hassani, Hamed
Ding, Dongsheng
Artificial Intelligence
Computation and Language
Machine Learning
Optimization and Control
The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safety. A promising approach is to enforce safety constraints through Reinforcement Learning from Human Feedback (RLHF). For such constrained RLHF, typical Lagrangian-based primal-dual policy optimization methods are computationally expensive and often unstable. This paper presents a perspective of dualization that reduces constrained alignment to an equivalent unconstrained alignment problem. We do so by pre-optimizing a smooth and convex dual function that has a closed form. This shortcut eliminates the need for cumbersome primal-dual policy iterations, greatly reducing the computational burden and improving training stability. Our strategy leads to two practical algorithms in model-based and preference-based settings (MoCAN and PeCAN, respectively). A broad range of experiments demonstrate the effectiveness and merits of our algorithms.
title One-Shot Safety Alignment for Large Language Models via Optimal Dualization
topic Artificial Intelligence
Computation and Language
Machine Learning
Optimization and Control
url https://arxiv.org/abs/2405.19544