The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Garcia, Basile, Qian, Crystal, Palminteri, Stefano
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910643630637056
author Garcia, Basile
Qian, Crystal
Palminteri, Stefano
author_facet Garcia, Basile
Qian, Crystal
Palminteri, Stefano
contents As large language models (LLMs) become increasingly integrated into society, their alignment with human morals is crucial. To better understand this alignment, we created a large corpus of human- and LLM-generated responses to various moral scenarios. We found a misalignment between human and LLM moral assessments; although both LLMs and humans tended to reject morally complex utilitarian dilemmas, LLMs were more sensitive to personal framing. We then conducted a quantitative user study involving 230 participants (N=230), who evaluated these responses by determining whether they were AI-generated and assessed their agreement with the responses. Human evaluators preferred LLMs' assessments in moral scenarios, though a systematic anti-AI bias was observed: participants were less likely to agree with judgments they believed to be machine-generated. Statistical and NLP-based analyses revealed subtle linguistic differences in responses, influencing detection and agreement. Overall, our findings highlight the complexities of human-AI perception in morally charged decision-making.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making
Garcia, Basile
Qian, Crystal
Palminteri, Stefano
Human-Computer Interaction
Artificial Intelligence
As large language models (LLMs) become increasingly integrated into society, their alignment with human morals is crucial. To better understand this alignment, we created a large corpus of human- and LLM-generated responses to various moral scenarios. We found a misalignment between human and LLM moral assessments; although both LLMs and humans tended to reject morally complex utilitarian dilemmas, LLMs were more sensitive to personal framing. We then conducted a quantitative user study involving 230 participants (N=230), who evaluated these responses by determining whether they were AI-generated and assessed their agreement with the responses. Human evaluators preferred LLMs' assessments in moral scenarios, though a systematic anti-AI bias was observed: participants were less likely to agree with judgments they believed to be machine-generated. Statistical and NLP-based analyses revealed subtle linguistic differences in responses, influencing detection and agreement. Overall, our findings highlight the complexities of human-AI perception in morally charged decision-making.
title The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2410.07304