Do LLMs have Consistent Values?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rozen, Naama, Bezalel, Liat, Elidan, Gal, Globerson, Amir, Daniel, Ella
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929542631784448
author Rozen, Naama
Bezalel, Liat
Elidan, Gal
Globerson, Amir
Daniel, Ella
author_facet Rozen, Naama
Bezalel, Liat
Elidan, Gal
Globerson, Amir
Daniel, Ella
contents Large Language Models (LLM) technology is constantly improving towards human-like dialogue. Values are a basic driving force underlying human behavior, but little research has been done to study the values exhibited in text generated by LLMs. Here we study this question by turning to the rich literature on value structure in psychology. We ask whether LLMs exhibit the same value structure that has been demonstrated in humans, including the ranking of values, and correlation between values. We show that the results of this analysis depend on how the LLM is prompted, and that under a particular prompting strategy (referred to as "Value Anchoring") the agreement with human data is quite compelling. Our results serve both to improve our understanding of values in LLMs, as well as introduce novel methods for assessing consistency in LLM responses.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12878
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do LLMs have Consistent Values?
Rozen, Naama
Bezalel, Liat
Elidan, Gal
Globerson, Amir
Daniel, Ella
Computation and Language
Artificial Intelligence
Large Language Models (LLM) technology is constantly improving towards human-like dialogue. Values are a basic driving force underlying human behavior, but little research has been done to study the values exhibited in text generated by LLMs. Here we study this question by turning to the rich literature on value structure in psychology. We ask whether LLMs exhibit the same value structure that has been demonstrated in humans, including the ranking of values, and correlation between values. We show that the results of this analysis depend on how the LLM is prompted, and that under a particular prompting strategy (referred to as "Value Anchoring") the agreement with human data is quite compelling. Our results serve both to improve our understanding of values in LLMs, as well as introduce novel methods for assessing consistency in LLM responses.
title Do LLMs have Consistent Values?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.12878