MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rosen, Simon, Singh, Siddarth, Gelo, Ebenezer, Robertson, Helen Sarah, Suder, Ibrahim, Williams, Victoria, Rosman, Benjamin, Tasse, Geraud Nangue, James, Steven
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911703972708352
author Rosen, Simon
Singh, Siddarth
Gelo, Ebenezer
Robertson, Helen Sarah
Suder, Ibrahim
Williams, Victoria
Rosman, Benjamin
Tasse, Geraud Nangue
James, Steven
author_facet Rosen, Simon
Singh, Siddarth
Gelo, Ebenezer
Robertson, Helen Sarah
Suder, Ibrahim
Williams, Victoria
Rosman, Benjamin
Tasse, Geraud Nangue
James, Steven
contents Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and cognitive science. We introduce Morality Chains, a novel formalism for representing moral norms as ordered deontic constraints, and MoralityGym, a benchmark of 98 ethical-dilemma problems presented as trolley-dilemma-style Gymnasium environments. By decoupling task-solving from moral evaluation and introducing a novel Morality Metric, MoralityGym allows the integration of insights from psychology and philosophy into the evaluation of norm-sensitive reasoning. Baseline results with Safe RL methods reveal key limitations, underscoring the need for more principled approaches to ethical decision-making. This work provides a foundation for developing AI systems that behave more reliably, transparently, and ethically in complex real-world contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13372
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
Rosen, Simon
Singh, Siddarth
Gelo, Ebenezer
Robertson, Helen Sarah
Suder, Ibrahim
Williams, Victoria
Rosman, Benjamin
Tasse, Geraud Nangue
James, Steven
Artificial Intelligence
Machine Learning
Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and cognitive science. We introduce Morality Chains, a novel formalism for representing moral norms as ordered deontic constraints, and MoralityGym, a benchmark of 98 ethical-dilemma problems presented as trolley-dilemma-style Gymnasium environments. By decoupling task-solving from moral evaluation and introducing a novel Morality Metric, MoralityGym allows the integration of insights from psychology and philosophy into the evaluation of norm-sensitive reasoning. Baseline results with Safe RL methods reveal key limitations, underscoring the need for more principled approaches to ethical decision-making. This work provides a foundation for developing AI systems that behave more reliably, transparently, and ethically in complex real-world contexts.
title MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.13372