Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Tianchen, Hairi, FNU, Yang, Haibo, Liu, Jia, Tong, Tian, Yang, Fan, Momma, Michinari, Gao, Yan
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2405.03082
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913345204912128
author Zhou, Tianchen
Hairi, FNU
Yang, Haibo
Liu, Jia
Tong, Tian
Yang, Fan
Momma, Michinari
Gao, Yan
author_facet Zhou, Tianchen
Hairi, FNU
Yang, Haibo
Liu, Jia
Tong, Tian
Yang, Fan
Momma, Michinari
Gao, Yan
contents Reinforcement learning with multiple, potentially conflicting objectives is pervasive in real-world applications, while this problem remains theoretically under-explored. This paper tackles the multi-objective reinforcement learning (MORL) problem and introduces an innovative actor-critic algorithm named MOAC which finds a policy by iteratively making trade-offs among conflicting reward signals. Notably, we provide the first analysis of finite-time Pareto-stationary convergence and corresponding sample complexity in both discounted and average reward settings. Our approach has two salient features: (a) MOAC mitigates the cumulative estimation bias resulting from finding an optimal common gradient descent direction out of stochastic samples. This enables provable convergence rate and sample complexity guarantees independent of the number of objectives; (b) With proper momentum coefficient, MOAC initializes the weights of individual policy gradients using samples from the environment, instead of manual initialization. This enhances the practicality and robustness of our algorithm. Finally, experiments conducted on a real-world dataset validate the effectiveness of our proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03082
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning
Zhou, Tianchen
Hairi, FNU
Yang, Haibo
Liu, Jia
Tong, Tian
Yang, Fan
Momma, Michinari
Gao, Yan
Machine Learning
Reinforcement learning with multiple, potentially conflicting objectives is pervasive in real-world applications, while this problem remains theoretically under-explored. This paper tackles the multi-objective reinforcement learning (MORL) problem and introduces an innovative actor-critic algorithm named MOAC which finds a policy by iteratively making trade-offs among conflicting reward signals. Notably, we provide the first analysis of finite-time Pareto-stationary convergence and corresponding sample complexity in both discounted and average reward settings. Our approach has two salient features: (a) MOAC mitigates the cumulative estimation bias resulting from finding an optimal common gradient descent direction out of stochastic samples. This enables provable convergence rate and sample complexity guarantees independent of the number of objectives; (b) With proper momentum coefficient, MOAC initializes the weights of individual policy gradients using samples from the environment, instead of manual initialization. This enhances the practicality and robustness of our algorithm. Finally, experiments conducted on a real-world dataset validate the effectiveness of our proposed method.
title Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2405.03082