Collective Constitutional AI: Aligning a Language Model with Public Input
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916283541356544 |
|---|---|
| author | Huang, Saffron Siddarth, Divya Lovitt, Liane Liao, Thomas I. Durmus, Esin Tamkin, Alex Ganguli, Deep |
| author_facet | Huang, Saffron Siddarth, Divya Lovitt, Liane Liao, Thomas I. Durmus, Esin Tamkin, Alex Ganguli, Deep |
| contents | There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To address this need, we present Collective Constitutional AI (CCAI): a multi-stage process for sourcing and integrating public input into LMs-from identifying a target population to sourcing principles to training and evaluating a model. We demonstrate the real-world practicality of this approach by creating what is, to our knowledge, the first LM fine-tuned with collectively sourced public input and evaluating this model against a baseline model trained with established principles from a LM developer. Our quantitative evaluations demonstrate several benefits of our approach: the CCAI-trained model shows lower bias across nine social dimensions compared to the baseline model, while maintaining equivalent performance on language, math, and helpful-harmless evaluations. Qualitative comparisons of the models suggest that the models differ on the basis of their respective constitutions, e.g., when prompted with contentious topics, the CCAI-trained model tends to generate responses that reframe the matter positively instead of a refusal. These results demonstrate a promising, tractable pathway toward publicly informed development of language models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_07814 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Collective Constitutional AI: Aligning a Language Model with Public Input Huang, Saffron Siddarth, Divya Lovitt, Liane Liao, Thomas I. Durmus, Esin Tamkin, Alex Ganguli, Deep Artificial Intelligence Computation and Language Human-Computer Interaction I.2.7; K.4.2 There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To address this need, we present Collective Constitutional AI (CCAI): a multi-stage process for sourcing and integrating public input into LMs-from identifying a target population to sourcing principles to training and evaluating a model. We demonstrate the real-world practicality of this approach by creating what is, to our knowledge, the first LM fine-tuned with collectively sourced public input and evaluating this model against a baseline model trained with established principles from a LM developer. Our quantitative evaluations demonstrate several benefits of our approach: the CCAI-trained model shows lower bias across nine social dimensions compared to the baseline model, while maintaining equivalent performance on language, math, and helpful-harmless evaluations. Qualitative comparisons of the models suggest that the models differ on the basis of their respective constitutions, e.g., when prompted with contentious topics, the CCAI-trained model tends to generate responses that reframe the matter positively instead of a refusal. These results demonstrate a promising, tractable pathway toward publicly informed development of language models. |
| title | Collective Constitutional AI: Aligning a Language Model with Public Input |
| topic | Artificial Intelligence Computation and Language Human-Computer Interaction I.2.7; K.4.2 |
| url | https://arxiv.org/abs/2406.07814 |