Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916742105661440 |
|---|---|
| author | Shakerinava, Mehran Ravanbakhsh, Siamak Oberman, Adam |
| author_facet | Shakerinava, Mehran Ravanbakhsh, Siamak Oberman, Adam |
| contents | Recent work has formalized the reward hypothesis through the lens of expected utility theory, by interpreting reward as utility. Hausner's foundational work showed that dropping the continuity axiom leads to a generalization of expected utility theory where utilities are lexicographically ordered vectors of arbitrary dimension. In this paper, we extend this result by identifying a simple and practical condition under which preferences cannot be represented by scalar rewards, necessitating a 2-dimensional reward function. We provide a full characterization of such reward functions, as well as the general d-dimensional case, in Markov Decision Processes (MDPs) under a memorylessness assumption on preferences. Furthermore, we show that optimal policies in this setting retain many desirable properties of their scalar-reward counterparts, while in the Constrained MDP (CMDP) setting -- another common multiobjective setting -- they do not. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_12049 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs Shakerinava, Mehran Ravanbakhsh, Siamak Oberman, Adam Machine Learning Artificial Intelligence Recent work has formalized the reward hypothesis through the lens of expected utility theory, by interpreting reward as utility. Hausner's foundational work showed that dropping the continuity axiom leads to a generalization of expected utility theory where utilities are lexicographically ordered vectors of arbitrary dimension. In this paper, we extend this result by identifying a simple and practical condition under which preferences cannot be represented by scalar rewards, necessitating a 2-dimensional reward function. We provide a full characterization of such reward functions, as well as the general d-dimensional case, in Markov Decision Processes (MDPs) under a memorylessness assumption on preferences. Furthermore, we show that optimal policies in this setting retain many desirable properties of their scalar-reward counterparts, while in the Constrained MDP (CMDP) setting -- another common multiobjective setting -- they do not. |
| title | Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2505.12049 |