Collective Constitutional AI: Aligning a Language Model with Public Input

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Saffron, Siddarth, Divya, Lovitt, Liane, Liao, Thomas I., Durmus, Esin, Tamkin, Alex, Ganguli, Deep
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916283541356544
author Huang, Saffron
Siddarth, Divya
Lovitt, Liane
Liao, Thomas I.
Durmus, Esin
Tamkin, Alex
Ganguli, Deep
author_facet Huang, Saffron
Siddarth, Divya
Lovitt, Liane
Liao, Thomas I.
Durmus, Esin
Tamkin, Alex
Ganguli, Deep
contents There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To address this need, we present Collective Constitutional AI (CCAI): a multi-stage process for sourcing and integrating public input into LMs-from identifying a target population to sourcing principles to training and evaluating a model. We demonstrate the real-world practicality of this approach by creating what is, to our knowledge, the first LM fine-tuned with collectively sourced public input and evaluating this model against a baseline model trained with established principles from a LM developer. Our quantitative evaluations demonstrate several benefits of our approach: the CCAI-trained model shows lower bias across nine social dimensions compared to the baseline model, while maintaining equivalent performance on language, math, and helpful-harmless evaluations. Qualitative comparisons of the models suggest that the models differ on the basis of their respective constitutions, e.g., when prompted with contentious topics, the CCAI-trained model tends to generate responses that reframe the matter positively instead of a refusal. These results demonstrate a promising, tractable pathway toward publicly informed development of language models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_07814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Collective Constitutional AI: Aligning a Language Model with Public Input
Huang, Saffron
Siddarth, Divya
Lovitt, Liane
Liao, Thomas I.
Durmus, Esin
Tamkin, Alex
Ganguli, Deep
Artificial Intelligence
Computation and Language
Human-Computer Interaction
I.2.7; K.4.2
There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To address this need, we present Collective Constitutional AI (CCAI): a multi-stage process for sourcing and integrating public input into LMs-from identifying a target population to sourcing principles to training and evaluating a model. We demonstrate the real-world practicality of this approach by creating what is, to our knowledge, the first LM fine-tuned with collectively sourced public input and evaluating this model against a baseline model trained with established principles from a LM developer. Our quantitative evaluations demonstrate several benefits of our approach: the CCAI-trained model shows lower bias across nine social dimensions compared to the baseline model, while maintaining equivalent performance on language, math, and helpful-harmless evaluations. Qualitative comparisons of the models suggest that the models differ on the basis of their respective constitutions, e.g., when prompted with contentious topics, the CCAI-trained model tends to generate responses that reframe the matter positively instead of a refusal. These results demonstrate a promising, tractable pathway toward publicly informed development of language models.
title Collective Constitutional AI: Aligning a Language Model with Public Input
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
I.2.7; K.4.2
url https://arxiv.org/abs/2406.07814