Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Allison, Pi, Yulu Niki, Mougan, Carlos
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910703488598016
author Huang, Allison
Pi, Yulu Niki
Mougan, Carlos
author_facet Huang, Allison
Pi, Yulu Niki
Mougan, Carlos
contents We explore how large language models (LLMs) can be influenced by prompting them to alter their initial decisions and align them with established ethical frameworks. Our study is based on two experiments designed to assess the susceptibility of LLMs to moral persuasion. In the first experiment, we examine the susceptibility to moral ambiguity by evaluating a Base Agent LLM on morally ambiguous scenarios and observing how a Persuader Agent attempts to modify the Base Agent's initial decisions. The second experiment evaluates the susceptibility of LLMs to align with predefined ethical frameworks by prompting them to adopt specific value alignments rooted in established philosophical theories. The results demonstrate that LLMs can indeed be persuaded in morally charged scenarios, with the success of persuasion depending on factors such as the model used, the complexity of the scenario, and the conversation length. Notably, LLMs of distinct sizes but from the same company produced markedly different outcomes, highlighting the variability in their susceptibility to ethical persuasion.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11731
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
Huang, Allison
Pi, Yulu Niki
Mougan, Carlos
Computation and Language
Artificial Intelligence
We explore how large language models (LLMs) can be influenced by prompting them to alter their initial decisions and align them with established ethical frameworks. Our study is based on two experiments designed to assess the susceptibility of LLMs to moral persuasion. In the first experiment, we examine the susceptibility to moral ambiguity by evaluating a Base Agent LLM on morally ambiguous scenarios and observing how a Persuader Agent attempts to modify the Base Agent's initial decisions. The second experiment evaluates the susceptibility of LLMs to align with predefined ethical frameworks by prompting them to adopt specific value alignments rooted in established philosophical theories. The results demonstrate that LLMs can indeed be persuaded in morally charged scenarios, with the success of persuasion depending on factors such as the model used, the complexity of the scenario, and the conversation length. Notably, LLMs of distinct sizes but from the same company produced markedly different outcomes, highlighting the variability in their susceptibility to ethical persuasion.
title Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.11731