More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Aaron J., Krishna, Satyapriya, Lakkaraju, Himabindu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913622307897344
author Li, Aaron J.
Krishna, Satyapriya
Lakkaraju, Himabindu
author_facet Li, Aaron J.
Krishna, Satyapriya
Lakkaraju, Himabindu
contents The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration alongside their cognitive performance. In practice, Reinforcement Learning From Human Feedback (RLHF) has been widely used to align LLMs with labeled human preferences, but its assumed effect on model trustworthiness hasn't been rigorously evaluated. To bridge this knowledge gap, this study investigates how models aligned with general-purpose preference data perform across five trustworthiness verticals: toxicity, stereotypical bias, machine ethics, truthfulness, and privacy. Our results demonstrate that RLHF on human preferences doesn't automatically guarantee trustworthiness, and reverse effects are often observed. Furthermore, we propose to adapt efficient influence function based data attribution methods to the RLHF setting to better understand the influence of fine-tuning data on individual trustworthiness benchmarks, and show its feasibility by providing our estimated attribution scores. Together, our results underscore the need for more nuanced approaches for model alignment from both the data and framework perspectives, and we hope this research will guide the community towards developing language models that are increasingly capable without sacrificing trustworthiness.
format Preprint
id arxiv_https___arxiv_org_abs_2404_18870
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
Li, Aaron J.
Krishna, Satyapriya
Lakkaraju, Himabindu
Computation and Language
Artificial Intelligence
The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration alongside their cognitive performance. In practice, Reinforcement Learning From Human Feedback (RLHF) has been widely used to align LLMs with labeled human preferences, but its assumed effect on model trustworthiness hasn't been rigorously evaluated. To bridge this knowledge gap, this study investigates how models aligned with general-purpose preference data perform across five trustworthiness verticals: toxicity, stereotypical bias, machine ethics, truthfulness, and privacy. Our results demonstrate that RLHF on human preferences doesn't automatically guarantee trustworthiness, and reverse effects are often observed. Furthermore, we propose to adapt efficient influence function based data attribution methods to the RLHF setting to better understand the influence of fine-tuning data on individual trustworthiness benchmarks, and show its feasibility by providing our estimated attribution scores. Together, our results underscore the need for more nuanced approaches for model alignment from both the data and framework perspectives, and we hope this research will guide the community towards developing language models that are increasingly capable without sacrificing trustworthiness.
title More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.18870