Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jain, Neel, Shrivastava, Aditya, Zhu, Chenyang, Liu, Daben, Samuel, Alfy, Panda, Ashwinee, Kumar, Anoop, Goldblum, Micah, Goldstein, Tom
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909758789779456
author Jain, Neel
Shrivastava, Aditya
Zhu, Chenyang
Liu, Daben
Samuel, Alfy
Panda, Ashwinee
Kumar, Anoop
Goldblum, Micah
Goldstein, Tom
author_facet Jain, Neel
Shrivastava, Aditya
Zhu, Chenyang
Liu, Daben
Samuel, Alfy
Panda, Ashwinee
Kumar, Anoop
Goldblum, Micah
Goldstein, Tom
contents A key component of building safe and reliable language models is enabling the models to appropriately refuse to follow certain instructions or answer certain questions. We may want models to output refusal messages for various categories of user queries, for example, ill-posed questions, instructions for committing illegal acts, or queries which require information past the model's knowledge horizon. Engineering models that refuse to answer such questions is complicated by the fact that an individual may want their model to exhibit varying levels of sensitivity for refusing queries of various categories, and different users may want different refusal rates. The current default approach involves training multiple models with varying proportions of refusal messages from each category to achieve the desired refusal rates, which is computationally expensive and may require training a new model to accommodate each user's desired preference over refusal rates. To address these challenges, we propose refusal tokens, one such token for each refusal category or a single refusal token, which are prepended to the model's responses during training. We then show how to increase or decrease the probability of generating the refusal token for each category during inference to steer the model's refusal behavior. Refusal tokens enable controlling a single model's refusal rates without the need of any further fine-tuning, but only by selectively intervening during generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06748
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
Jain, Neel
Shrivastava, Aditya
Zhu, Chenyang
Liu, Daben
Samuel, Alfy
Panda, Ashwinee
Kumar, Anoop
Goldblum, Micah
Goldstein, Tom
Machine Learning
Computation and Language
A key component of building safe and reliable language models is enabling the models to appropriately refuse to follow certain instructions or answer certain questions. We may want models to output refusal messages for various categories of user queries, for example, ill-posed questions, instructions for committing illegal acts, or queries which require information past the model's knowledge horizon. Engineering models that refuse to answer such questions is complicated by the fact that an individual may want their model to exhibit varying levels of sensitivity for refusing queries of various categories, and different users may want different refusal rates. The current default approach involves training multiple models with varying proportions of refusal messages from each category to achieve the desired refusal rates, which is computationally expensive and may require training a new model to accommodate each user's desired preference over refusal rates. To address these challenges, we propose refusal tokens, one such token for each refusal category or a single refusal token, which are prepended to the model's responses during training. We then show how to increase or decrease the probability of generating the refusal token for each category during inference to steer the model's refusal behavior. Refusal tokens enable controlling a single model's refusal rates without the need of any further fine-tuning, but only by selectively intervening during generation.
title Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2412.06748