Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Obi, Ike, Pant, Rohan, Agrawal, Srishti Shekhar, Ghazanfar, Maham, Basiletti, Aaron
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912123907473408
author Obi, Ike
Pant, Rohan
Agrawal, Srishti Shekhar
Ghazanfar, Maham
Basiletti, Aaron
author_facet Obi, Ike
Pant, Rohan
Agrawal, Srishti Shekhar
Ghazanfar, Maham
Basiletti, Aaron
contents LLMs are increasingly fine-tuned using RLHF datasets to align them with human preferences and values. However, very limited research has investigated which specific human values are operationalized through these datasets. In this paper, we introduce Value Imprint, a framework for auditing and classifying the human values embedded within RLHF datasets. To investigate the viability of this framework, we conducted three case study experiments by auditing the Anthropic/hh-rlhf, OpenAI WebGPT Comparisons, and Alpaca GPT-4-LLM datasets to examine the human values embedded within them. Our analysis involved a two-phase process. During the first phase, we developed a taxonomy of human values through an integrated review of prior works from philosophy, axiology, and ethics. Then, we applied this taxonomy to annotate 6,501 RLHF preferences. During the second phase, we employed the labels generated from the annotation as ground truth data for training a transformer-based machine learning model to audit and classify the three RLHF datasets. Through this approach, we discovered that information-utility values, including Wisdom/Knowledge and Information Seeking, were the most dominant human values within all three RLHF datasets. In contrast, prosocial and democratic values, including Well-being, Justice, and Human/Animal Rights, were the least represented human values. These findings have significant implications for developing language models that align with societal values and norms. We contribute our datasets to support further research in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11937
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets
Obi, Ike
Pant, Rohan
Agrawal, Srishti Shekhar
Ghazanfar, Maham
Basiletti, Aaron
Machine Learning
Artificial Intelligence
LLMs are increasingly fine-tuned using RLHF datasets to align them with human preferences and values. However, very limited research has investigated which specific human values are operationalized through these datasets. In this paper, we introduce Value Imprint, a framework for auditing and classifying the human values embedded within RLHF datasets. To investigate the viability of this framework, we conducted three case study experiments by auditing the Anthropic/hh-rlhf, OpenAI WebGPT Comparisons, and Alpaca GPT-4-LLM datasets to examine the human values embedded within them. Our analysis involved a two-phase process. During the first phase, we developed a taxonomy of human values through an integrated review of prior works from philosophy, axiology, and ethics. Then, we applied this taxonomy to annotate 6,501 RLHF preferences. During the second phase, we employed the labels generated from the annotation as ground truth data for training a transformer-based machine learning model to audit and classify the three RLHF datasets. Through this approach, we discovered that information-utility values, including Wisdom/Knowledge and Information Seeking, were the most dominant human values within all three RLHF datasets. In contrast, prosocial and democratic values, including Well-being, Justice, and Human/Animal Rights, were the least represented human values. These findings have significant implications for developing language models that align with societal values and norms. We contribute our datasets to support further research in this area.
title Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2411.11937