Absolute Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Weiye, Li, Feihan, Sun, Yifan, Chen, Rui, Wei, Tianhao, Liu, Changliu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916265470197760
author Zhao, Weiye
Li, Feihan
Sun, Yifan
Chen, Rui
Wei, Tianhao
Liu, Changliu
author_facet Zhao, Weiye
Li, Feihan
Sun, Yifan
Chen, Rui
Wei, Tianhao
Liu, Changliu
contents In recent years, trust region on-policy reinforcement learning has achieved impressive results in addressing complex control tasks and gaming scenarios. However, contemporary state-of-the-art algorithms within this category primarily emphasize improvement in expected performance, lacking the ability to control over the worst-case performance outcomes. To address this limitation, we introduce a novel objective function, optimizing which leads to guaranteed monotonic improvement in the lower probability bound of performance with high confidence. Building upon this groundbreaking theoretical advancement, we further introduce a practical solution called Absolute Policy Optimization (APO). Our experiments demonstrate the effectiveness of our approach across challenging continuous control benchmark tasks and extend its applicability to mastering Atari games. Our findings reveal that APO as well as its efficient variation Proximal Absolute Policy Optimization (PAPO) significantly outperforms state-of-the-art policy gradient algorithms, resulting in substantial improvements in worst-case performance, as well as expected performance.
format Preprint
id arxiv_https___arxiv_org_abs_2310_13230
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Absolute Policy Optimization
Zhao, Weiye
Li, Feihan
Sun, Yifan
Chen, Rui
Wei, Tianhao
Liu, Changliu
Machine Learning
Artificial Intelligence
Robotics
In recent years, trust region on-policy reinforcement learning has achieved impressive results in addressing complex control tasks and gaming scenarios. However, contemporary state-of-the-art algorithms within this category primarily emphasize improvement in expected performance, lacking the ability to control over the worst-case performance outcomes. To address this limitation, we introduce a novel objective function, optimizing which leads to guaranteed monotonic improvement in the lower probability bound of performance with high confidence. Building upon this groundbreaking theoretical advancement, we further introduce a practical solution called Absolute Policy Optimization (APO). Our experiments demonstrate the effectiveness of our approach across challenging continuous control benchmark tasks and extend its applicability to mastering Atari games. Our findings reveal that APO as well as its efficient variation Proximal Absolute Policy Optimization (PAPO) significantly outperforms state-of-the-art policy gradient algorithms, resulting in substantial improvements in worst-case performance, as well as expected performance.
title Absolute Policy Optimization
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2310.13230