An Online Multiobjective Policy Gradient for Long-run Average-reward Markov Decision Process

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Misra, Rahul, Bujorianu, Manuela L., Wisniewski, Rafał
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914161020108800
author Misra, Rahul
Bujorianu, Manuela L.
Wisniewski, Rafał
author_facet Misra, Rahul
Bujorianu, Manuela L.
Wisniewski, Rafał
contents We propose a reinforcement learning (RL) framework for multi-objective decision-making, where the agent seeks to optimize a vector of rewards rather than a single scalar value. The objective is to ensure that the time-averaged reward vector converges asymptotically to a predefined target set. Since standard RL algorithms operate on scalar rewards, we introduce a dynamic scalarization mechanism guided by Blackwell's Approachability Theorem. This theorem enables adaptive updates of the scalarization vector to guarantee convergence toward the target set. Assuming ergodicity, the Markov chain induced by the learned policies admits a stationary distribution, ensuring all states recur with finite return times. Our algorithm exploits this property by defining an inner loop that applies a policy gradient method (with baseline) between successive visits to a designated recurrent state, enforcing Blackwell's condition at each iteration. An outer loop then updates the scalarization vector after each recurrence. We establish theoretical convergence of the long-run average reward vector to the target set and validate the approach through a numerical example.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Online Multiobjective Policy Gradient for Long-run Average-reward Markov Decision Process
Misra, Rahul
Bujorianu, Manuela L.
Wisniewski, Rafał
Systems and Control
Optimization and Control
Probability
We propose a reinforcement learning (RL) framework for multi-objective decision-making, where the agent seeks to optimize a vector of rewards rather than a single scalar value. The objective is to ensure that the time-averaged reward vector converges asymptotically to a predefined target set. Since standard RL algorithms operate on scalar rewards, we introduce a dynamic scalarization mechanism guided by Blackwell's Approachability Theorem. This theorem enables adaptive updates of the scalarization vector to guarantee convergence toward the target set. Assuming ergodicity, the Markov chain induced by the learned policies admits a stationary distribution, ensuring all states recur with finite return times. Our algorithm exploits this property by defining an inner loop that applies a policy gradient method (with baseline) between successive visits to a designated recurrent state, enforcing Blackwell's condition at each iteration. An outer loop then updates the scalarization vector after each recurrence. We establish theoretical convergence of the long-run average reward vector to the target set and validate the approach through a numerical example.
title An Online Multiobjective Policy Gradient for Long-run Average-reward Markov Decision Process
topic Systems and Control
Optimization and Control
Probability
url https://arxiv.org/abs/2511.13034