Deliberative Alignment: Reasoning Enables Safer Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913642322067456 |
|---|---|
| author | Guan, Melody Y. Joglekar, Manas Wallace, Eric Jain, Saachi Barak, Boaz Helyar, Alec Dias, Rachel Vallone, Andrea Ren, Hongyu Wei, Jason Chung, Hyung Won Toyer, Sam Heidecke, Johannes Beutel, Alex Glaese, Amelia |
| author_facet | Guan, Melody Y. Joglekar, Manas Wallace, Eric Jain, Saachi Barak, Boaz Helyar, Alec Dias, Rachel Vallone, Andrea Ren, Hongyu Wei, Jason Chung, Hyung Won Toyer, Sam Heidecke, Johannes Beutel, Alex Glaese, Amelia |
| contents | As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly teaches the model safety specifications and trains it to explicitly recall and accurately reason over the specifications before answering. We used this approach to align OpenAI's o-series models, and achieved highly precise adherence to OpenAI's safety policies, without requiring human-written chain-of-thoughts or answers. Deliberative Alignment pushes the Pareto frontier by simultaneously increasing robustness to jailbreaks while decreasing overrefusal rates, and also improves out-of-distribution generalization. We demonstrate that reasoning over explicitly specified policies enables more scalable, trustworthy, and interpretable alignment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_16339 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Deliberative Alignment: Reasoning Enables Safer Language Models Guan, Melody Y. Joglekar, Manas Wallace, Eric Jain, Saachi Barak, Boaz Helyar, Alec Dias, Rachel Vallone, Andrea Ren, Hongyu Wei, Jason Chung, Hyung Won Toyer, Sam Heidecke, Johannes Beutel, Alex Glaese, Amelia Computation and Language Artificial Intelligence Computers and Society Machine Learning As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly teaches the model safety specifications and trains it to explicitly recall and accurately reason over the specifications before answering. We used this approach to align OpenAI's o-series models, and achieved highly precise adherence to OpenAI's safety policies, without requiring human-written chain-of-thoughts or answers. Deliberative Alignment pushes the Pareto frontier by simultaneously increasing robustness to jailbreaks while decreasing overrefusal rates, and also improves out-of-distribution generalization. We demonstrate that reasoning over explicitly specified policies enables more scalable, trustworthy, and interpretable alignment. |
| title | Deliberative Alignment: Reasoning Enables Safer Language Models |
| topic | Computation and Language Artificial Intelligence Computers and Society Machine Learning |
| url | https://arxiv.org/abs/2412.16339 |