Suppressing Pink Elephants with Direct Principle Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Castricato, Louis, Lile, Nathan, Anand, Suraj, Schoelkopf, Hailey, Verma, Siddharth, Biderman, Stella
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909104840114176
author Castricato, Louis
Lile, Nathan
Anand, Suraj
Schoelkopf, Hailey
Verma, Siddharth
Biderman, Stella
author_facet Castricato, Louis
Lile, Nathan
Anand, Suraj
Schoelkopf, Hailey
Verma, Siddharth
Biderman, Stella
contents Existing methods for controlling language models, such as RLHF and Constitutional AI, involve determining which LLM behaviors are desirable and training them into a language model. However, in many cases, it is desirable for LLMs to be controllable at inference time, so that they can be used in multiple contexts with diverse needs. We illustrate this with the Pink Elephant Problem: instructing an LLM to avoid discussing a certain entity (a ``Pink Elephant''), and instead discuss a preferred entity (``Grey Elephant''). We apply a novel simplification of Constitutional AI, Direct Principle Feedback, which skips the ranking of responses and uses DPO directly on critiques and revisions. Our results show that after DPF fine-tuning on our synthetic Pink Elephants dataset, our 13B fine-tuned LLaMA 2 model significantly outperforms Llama-2-13B-Chat and a prompted baseline, and performs as well as GPT-4 in on our curated test set assessing the Pink Elephant Problem.
format Preprint
id arxiv_https___arxiv_org_abs_2402_07896
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Suppressing Pink Elephants with Direct Principle Feedback
Castricato, Louis
Lile, Nathan
Anand, Suraj
Schoelkopf, Hailey
Verma, Siddharth
Biderman, Stella
Computation and Language
Existing methods for controlling language models, such as RLHF and Constitutional AI, involve determining which LLM behaviors are desirable and training them into a language model. However, in many cases, it is desirable for LLMs to be controllable at inference time, so that they can be used in multiple contexts with diverse needs. We illustrate this with the Pink Elephant Problem: instructing an LLM to avoid discussing a certain entity (a ``Pink Elephant''), and instead discuss a preferred entity (``Grey Elephant''). We apply a novel simplification of Constitutional AI, Direct Principle Feedback, which skips the ranking of responses and uses DPO directly on critiques and revisions. Our results show that after DPF fine-tuning on our synthetic Pink Elephants dataset, our 13B fine-tuned LLaMA 2 model significantly outperforms Llama-2-13B-Chat and a prompted baseline, and performs as well as GPT-4 in on our curated test set assessing the Pink Elephant Problem.
title Suppressing Pink Elephants with Direct Principle Feedback
topic Computation and Language
url https://arxiv.org/abs/2402.07896