Thresholded Lexicographic Ordered Multiobjective Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tercan, Alperen, Prabhu, Vinayak S.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910588561522688
author Tercan, Alperen
Prabhu, Vinayak S.
author_facet Tercan, Alperen
Prabhu, Vinayak S.
contents Lexicographic multi-objective problems, which impose a lexicographic importance order over the objectives, arise in many real-life scenarios. Existing Reinforcement Learning work directly addressing lexicographic tasks has been scarce. The few proposed approaches were all noted to be heuristics without theoretical guarantees as the Bellman equation is not applicable to them. Additionally, the practical applicability of these prior approaches also suffers from various issues such as not being able to reach the goal state. While some of these issues have been known before, in this work we investigate further shortcomings, and propose fixes for improving practical performance in many cases. We also present a policy optimization approach using our Lexicographic Projection Optimization (LPO) algorithm that has the potential to address these theoretical and practical concerns. Finally, we demonstrate our proposed algorithms on benchmark problems.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13493
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Thresholded Lexicographic Ordered Multiobjective Reinforcement Learning
Tercan, Alperen
Prabhu, Vinayak S.
Machine Learning
Artificial Intelligence
Lexicographic multi-objective problems, which impose a lexicographic importance order over the objectives, arise in many real-life scenarios. Existing Reinforcement Learning work directly addressing lexicographic tasks has been scarce. The few proposed approaches were all noted to be heuristics without theoretical guarantees as the Bellman equation is not applicable to them. Additionally, the practical applicability of these prior approaches also suffers from various issues such as not being able to reach the goal state. While some of these issues have been known before, in this work we investigate further shortcomings, and propose fixes for improving practical performance in many cases. We also present a policy optimization approach using our Lexicographic Projection Optimization (LPO) algorithm that has the potential to address these theoretical and practical concerns. Finally, we demonstrate our proposed algorithms on benchmark problems.
title Thresholded Lexicographic Ordered Multiobjective Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2408.13493