Are Language Models Consequentialist or Deontological Moral Reasoners?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Samway, Keenan, Kleiman-Weiner, Max, Piedrahita, David Guzman, Mihalcea, Rada, Schölkopf, Bernhard, Jin, Zhijing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912643768385536
author Samway, Keenan
Kleiman-Weiner, Max
Piedrahita, David Guzman
Mihalcea, Rada
Schölkopf, Bernhard
Jin, Zhijing
author_facet Samway, Keenan
Kleiman-Weiner, Max
Piedrahita, David Guzman
Mihalcea, Rada
Schölkopf, Bernhard
Jin, Zhijing
contents As AI systems increasingly navigate applications in healthcare, law, and governance, understanding how they handle ethically complex scenarios becomes critical. Previous work has mainly examined the moral judgments in large language models (LLMs), rather than their underlying moral reasoning process. In contrast, we focus on a large-scale analysis of the moral reasoning traces provided by LLMs. Furthermore, unlike prior work that attempted to draw inferences from only a handful of moral dilemmas, our study leverages over 600 distinct trolley problems as probes for revealing the reasoning patterns that emerge within different LLMs. We introduce and test a taxonomy of moral rationales to systematically classify reasoning traces according to two main normative ethical theories: consequentialism and deontology. Our analysis reveals that LLM chains-of-thought tend to favor deontological principles based on moral obligations, while post-hoc explanations shift notably toward consequentialist rationales that emphasize utility. Our framework provides a foundation for understanding how LLMs process and articulate ethical considerations, an important step toward safe and interpretable deployment of LLMs in high-stakes decision-making environments. Our code is available at https://github.com/keenansamway/moral-lens .
format Preprint
id arxiv_https___arxiv_org_abs_2505_21479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Are Language Models Consequentialist or Deontological Moral Reasoners?
Samway, Keenan
Kleiman-Weiner, Max
Piedrahita, David Guzman
Mihalcea, Rada
Schölkopf, Bernhard
Jin, Zhijing
Computation and Language
As AI systems increasingly navigate applications in healthcare, law, and governance, understanding how they handle ethically complex scenarios becomes critical. Previous work has mainly examined the moral judgments in large language models (LLMs), rather than their underlying moral reasoning process. In contrast, we focus on a large-scale analysis of the moral reasoning traces provided by LLMs. Furthermore, unlike prior work that attempted to draw inferences from only a handful of moral dilemmas, our study leverages over 600 distinct trolley problems as probes for revealing the reasoning patterns that emerge within different LLMs. We introduce and test a taxonomy of moral rationales to systematically classify reasoning traces according to two main normative ethical theories: consequentialism and deontology. Our analysis reveals that LLM chains-of-thought tend to favor deontological principles based on moral obligations, while post-hoc explanations shift notably toward consequentialist rationales that emphasize utility. Our framework provides a foundation for understanding how LLMs process and articulate ethical considerations, an important step toward safe and interpretable deployment of LLMs in high-stakes decision-making environments. Our code is available at https://github.com/keenansamway/moral-lens .
title Are Language Models Consequentialist or Deontological Moral Reasoners?
topic Computation and Language
url https://arxiv.org/abs/2505.21479