One-Shot Safety Alignment for Large Language Models via Optimal Dualization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913583020900352 |
|---|---|
| author | Huang, Xinmeng Li, Shuo Dobriban, Edgar Bastani, Osbert Hassani, Hamed Ding, Dongsheng |
| author_facet | Huang, Xinmeng Li, Shuo Dobriban, Edgar Bastani, Osbert Hassani, Hamed Ding, Dongsheng |
| contents | The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safety. A promising approach is to enforce safety constraints through Reinforcement Learning from Human Feedback (RLHF). For such constrained RLHF, typical Lagrangian-based primal-dual policy optimization methods are computationally expensive and often unstable. This paper presents a perspective of dualization that reduces constrained alignment to an equivalent unconstrained alignment problem. We do so by pre-optimizing a smooth and convex dual function that has a closed form. This shortcut eliminates the need for cumbersome primal-dual policy iterations, greatly reducing the computational burden and improving training stability. Our strategy leads to two practical algorithms in model-based and preference-based settings (MoCAN and PeCAN, respectively). A broad range of experiments demonstrate the effectiveness and merits of our algorithms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_19544 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | One-Shot Safety Alignment for Large Language Models via Optimal Dualization Huang, Xinmeng Li, Shuo Dobriban, Edgar Bastani, Osbert Hassani, Hamed Ding, Dongsheng Artificial Intelligence Computation and Language Machine Learning Optimization and Control The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safety. A promising approach is to enforce safety constraints through Reinforcement Learning from Human Feedback (RLHF). For such constrained RLHF, typical Lagrangian-based primal-dual policy optimization methods are computationally expensive and often unstable. This paper presents a perspective of dualization that reduces constrained alignment to an equivalent unconstrained alignment problem. We do so by pre-optimizing a smooth and convex dual function that has a closed form. This shortcut eliminates the need for cumbersome primal-dual policy iterations, greatly reducing the computational burden and improving training stability. Our strategy leads to two practical algorithms in model-based and preference-based settings (MoCAN and PeCAN, respectively). A broad range of experiments demonstrate the effectiveness and merits of our algorithms. |
| title | One-Shot Safety Alignment for Large Language Models via Optimal Dualization |
| topic | Artificial Intelligence Computation and Language Machine Learning Optimization and Control |
| url | https://arxiv.org/abs/2405.19544 |