Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shakerinava, Mehran, Ravanbakhsh, Siamak, Oberman, Adam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916742105661440
author Shakerinava, Mehran
Ravanbakhsh, Siamak
Oberman, Adam
author_facet Shakerinava, Mehran
Ravanbakhsh, Siamak
Oberman, Adam
contents Recent work has formalized the reward hypothesis through the lens of expected utility theory, by interpreting reward as utility. Hausner's foundational work showed that dropping the continuity axiom leads to a generalization of expected utility theory where utilities are lexicographically ordered vectors of arbitrary dimension. In this paper, we extend this result by identifying a simple and practical condition under which preferences cannot be represented by scalar rewards, necessitating a 2-dimensional reward function. We provide a full characterization of such reward functions, as well as the general d-dimensional case, in Markov Decision Processes (MDPs) under a memorylessness assumption on preferences. Furthermore, we show that optimal policies in this setting retain many desirable properties of their scalar-reward counterparts, while in the Constrained MDP (CMDP) setting -- another common multiobjective setting -- they do not.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
Shakerinava, Mehran
Ravanbakhsh, Siamak
Oberman, Adam
Machine Learning
Artificial Intelligence
Recent work has formalized the reward hypothesis through the lens of expected utility theory, by interpreting reward as utility. Hausner's foundational work showed that dropping the continuity axiom leads to a generalization of expected utility theory where utilities are lexicographically ordered vectors of arbitrary dimension. In this paper, we extend this result by identifying a simple and practical condition under which preferences cannot be represented by scalar rewards, necessitating a 2-dimensional reward function. We provide a full characterization of such reward functions, as well as the general d-dimensional case, in Markov Decision Processes (MDPs) under a memorylessness assumption on preferences. Furthermore, we show that optimal policies in this setting retain many desirable properties of their scalar-reward counterparts, while in the Constrained MDP (CMDP) setting -- another common multiobjective setting -- they do not.
title Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.12049