Inertia in Moral and Value Judgments of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Bruce W., Lee, Yeongheon, Cho, Hyunsoo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913043202441216
author Lee, Bruce W.
Lee, Yeongheon
Cho, Hyunsoo
author_facet Lee, Bruce W.
Lee, Yeongheon
Cho, Hyunsoo
contents Large Language Models (LLMs) behave non-deterministically, and prompting has become a common method for steering their outputs. A popular strategy is to assign a persona to the model to produce more varied, context-sensitive responses, similar to how responses vary across human individuals. Against the expectation that persona prompting yields a wide range of opinions, our experiments show that LLMs keep consistent value orientations. We observe a persistent inertia in their responses, where certain moral and value dimensions (especially harm avoidance and fairness) stay skewed in one direction across persona settings. To study this, we use role-play at scale, which pairs randomized persona prompts with a macro-level analysis of model outputs. Our results point to strong internal biases and value preferences in LLMs, which we call value orientation and inertia. These models warrant scrutiny and adjustment before use in applications where balanced outputs matter.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Inertia in Moral and Value Judgments of Large Language Models
Lee, Bruce W.
Lee, Yeongheon
Cho, Hyunsoo
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Large Language Models (LLMs) behave non-deterministically, and prompting has become a common method for steering their outputs. A popular strategy is to assign a persona to the model to produce more varied, context-sensitive responses, similar to how responses vary across human individuals. Against the expectation that persona prompting yields a wide range of opinions, our experiments show that LLMs keep consistent value orientations. We observe a persistent inertia in their responses, where certain moral and value dimensions (especially harm avoidance and fairness) stay skewed in one direction across persona settings. To study this, we use role-play at scale, which pairs randomized persona prompts with a macro-level analysis of model outputs. Our results point to strong internal biases and value preferences in LLMs, which we call value orientation and inertia. These models warrant scrutiny and adjustment before use in applications where balanced outputs matter.
title Inertia in Moral and Value Judgments of Large Language Models
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2408.09049