Deliberative Alignment: Reasoning Enables Safer Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Melody Y., Joglekar, Manas, Wallace, Eric, Jain, Saachi, Barak, Boaz, Helyar, Alec, Dias, Rachel, Vallone, Andrea, Ren, Hongyu, Wei, Jason, Chung, Hyung Won, Toyer, Sam, Heidecke, Johannes, Beutel, Alex, Glaese, Amelia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913642322067456
author Guan, Melody Y.
Joglekar, Manas
Wallace, Eric
Jain, Saachi
Barak, Boaz
Helyar, Alec
Dias, Rachel
Vallone, Andrea
Ren, Hongyu
Wei, Jason
Chung, Hyung Won
Toyer, Sam
Heidecke, Johannes
Beutel, Alex
Glaese, Amelia
author_facet Guan, Melody Y.
Joglekar, Manas
Wallace, Eric
Jain, Saachi
Barak, Boaz
Helyar, Alec
Dias, Rachel
Vallone, Andrea
Ren, Hongyu
Wei, Jason
Chung, Hyung Won
Toyer, Sam
Heidecke, Johannes
Beutel, Alex
Glaese, Amelia
contents As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly teaches the model safety specifications and trains it to explicitly recall and accurately reason over the specifications before answering. We used this approach to align OpenAI's o-series models, and achieved highly precise adherence to OpenAI's safety policies, without requiring human-written chain-of-thoughts or answers. Deliberative Alignment pushes the Pareto frontier by simultaneously increasing robustness to jailbreaks while decreasing overrefusal rates, and also improves out-of-distribution generalization. We demonstrate that reasoning over explicitly specified policies enables more scalable, trustworthy, and interpretable alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16339
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deliberative Alignment: Reasoning Enables Safer Language Models
Guan, Melody Y.
Joglekar, Manas
Wallace, Eric
Jain, Saachi
Barak, Boaz
Helyar, Alec
Dias, Rachel
Vallone, Andrea
Ren, Hongyu
Wei, Jason
Chung, Hyung Won
Toyer, Sam
Heidecke, Johannes
Beutel, Alex
Glaese, Amelia
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly teaches the model safety specifications and trains it to explicitly recall and accurately reason over the specifications before answering. We used this approach to align OpenAI's o-series models, and achieved highly precise adherence to OpenAI's safety policies, without requiring human-written chain-of-thoughts or answers. Deliberative Alignment pushes the Pareto frontier by simultaneously increasing robustness to jailbreaks while decreasing overrefusal rates, and also improves out-of-distribution generalization. We demonstrate that reasoning over explicitly specified policies enables more scalable, trustworthy, and interpretable alignment.
title Deliberative Alignment: Reasoning Enables Safer Language Models
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2412.16339