Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fränken, Jan-Philipp, Zelikman, Eric, Rafailov, Rafael, Gandhi, Kanishk, Gerstenberg, Tobias, Goodman, Noah D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914803737427968
author Fränken, Jan-Philipp
Zelikman, Eric
Rafailov, Rafael
Gandhi, Kanishk
Gerstenberg, Tobias
Goodman, Noah D.
author_facet Fränken, Jan-Philipp
Zelikman, Eric
Rafailov, Rafael
Gandhi, Kanishk
Gerstenberg, Tobias
Goodman, Noah D.
contents When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles (i.e., a constitution) into a model is resource-intensive, technically challenging, and generally requires human preference labels or examples. We introduce SAMI, an iterative algorithm that finetunes a pretrained language model (without requiring preference labels or demonstrations) to increase the conditional mutual information between constitutions and self-generated responses given queries from a dataset. On single-turn dialogue and summarization, a SAMI-trained mistral-7b outperforms the initial pretrained model, with win rates between 66% and 77%. Strikingly, it also surpasses an instruction-finetuned baseline (mistral-7b-instruct) with win rates between 55% and 57% on single-turn dialogue. SAMI requires a model that writes the principles. To avoid dependence on strong models for writing principles, we align a strong pretrained model (mixtral-8x7b) using constitutions written by a weak instruction-finetuned model (mistral-7b-instruct), achieving a 65% win rate on summarization. Finally, we investigate whether SAMI generalizes to diverse summarization principles (e.g., "summaries should be scientific") and scales to stronger models (llama3-70b), finding that it achieves win rates of up to 68% for learned and 67% for held-out principles compared to the base model. Our results show that a pretrained LM can learn to follow constitutions without using preference labels, demonstrations, or human oversight.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
Fränken, Jan-Philipp
Zelikman, Eric
Rafailov, Rafael
Gandhi, Kanishk
Gerstenberg, Tobias
Goodman, Noah D.
Computation and Language
When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles (i.e., a constitution) into a model is resource-intensive, technically challenging, and generally requires human preference labels or examples. We introduce SAMI, an iterative algorithm that finetunes a pretrained language model (without requiring preference labels or demonstrations) to increase the conditional mutual information between constitutions and self-generated responses given queries from a dataset. On single-turn dialogue and summarization, a SAMI-trained mistral-7b outperforms the initial pretrained model, with win rates between 66% and 77%. Strikingly, it also surpasses an instruction-finetuned baseline (mistral-7b-instruct) with win rates between 55% and 57% on single-turn dialogue. SAMI requires a model that writes the principles. To avoid dependence on strong models for writing principles, we align a strong pretrained model (mixtral-8x7b) using constitutions written by a weak instruction-finetuned model (mistral-7b-instruct), achieving a 65% win rate on summarization. Finally, we investigate whether SAMI generalizes to diverse summarization principles (e.g., "summaries should be scientific") and scales to stronger models (llama3-70b), finding that it achieves win rates of up to 68% for learned and 67% for held-out principles compared to the base model. Our results show that a pretrained LM can learn to follow constitutions without using preference labels, demonstrations, or human oversight.
title Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
topic Computation and Language
url https://arxiv.org/abs/2404.14313