Addressing Moral Uncertainty using Large Language Models for Ethical Decision-Making

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dubey, Rohit K., Dailisan, Damian, Mahajan, Sachit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910188835962880
author Dubey, Rohit K.
Dailisan, Damian
Mahajan, Sachit
author_facet Dubey, Rohit K.
Dailisan, Damian
Mahajan, Sachit
contents We present an ethical decision-making framework that refines a pre-trained reinforcement learning (RL) model using a task-agnostic ethical layer. Following initial training, the RL model undergoes ethical fine-tuning, where human feedback is replaced by feedback generated from a large language model (LLM). The LLM embodies consequentialist, deontological, virtue, social justice, and care ethics as moral principles to assign belief values to recommended actions during ethical decision-making. An ethical layer aggregates belief scores from multiple LLM-derived moral perspectives using Belief Jensen-Shannon Divergence and Dempster-Shafer Theory into probability scores that also serve as the shaping reward, steering the agent toward choices that align with a balanced ethical framework. This integrated learning framework helps the RL agent navigate moral uncertainty in complex environments and enables it to make morally sound decisions across diverse tasks. Our approach, tested across different LLM variants and compared with other belief aggregation techniques, demonstrates improved consistency, adaptability, and reduced reliance on handcrafted ethical rewards. This method is especially effective in dynamic scenarios where ethical challenges arise unexpectedly, making it well-suited for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05724
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Addressing Moral Uncertainty using Large Language Models for Ethical Decision-Making
Dubey, Rohit K.
Dailisan, Damian
Mahajan, Sachit
Computers and Society
Artificial Intelligence
We present an ethical decision-making framework that refines a pre-trained reinforcement learning (RL) model using a task-agnostic ethical layer. Following initial training, the RL model undergoes ethical fine-tuning, where human feedback is replaced by feedback generated from a large language model (LLM). The LLM embodies consequentialist, deontological, virtue, social justice, and care ethics as moral principles to assign belief values to recommended actions during ethical decision-making. An ethical layer aggregates belief scores from multiple LLM-derived moral perspectives using Belief Jensen-Shannon Divergence and Dempster-Shafer Theory into probability scores that also serve as the shaping reward, steering the agent toward choices that align with a balanced ethical framework. This integrated learning framework helps the RL agent navigate moral uncertainty in complex environments and enables it to make morally sound decisions across diverse tasks. Our approach, tested across different LLM variants and compared with other belief aggregation techniques, demonstrates improved consistency, adaptability, and reduced reliance on handcrafted ethical rewards. This method is especially effective in dynamic scenarios where ethical challenges arise unexpectedly, making it well-suited for real-world applications.
title Addressing Moral Uncertainty using Large Language Models for Ethical Decision-Making
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2503.05724