Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Yixiu, Wang, Qi, Chen, Chen, Qu, Yun, Ji, Xiangyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912099401203712
author Mao, Yixiu
Wang, Qi
Chen, Chen
Qu, Yun
Ji, Xiangyang
author_facet Mao, Yixiu
Wang, Qi
Chen, Chen
Qu, Yun
Ji, Xiangyang
contents In offline reinforcement learning (RL), addressing the out-of-distribution (OOD) action issue has been a focus, but we argue that there exists an OOD state issue that also impairs performance yet has been underexplored. Such an issue describes the scenario when the agent encounters states out of the offline dataset during the test phase, leading to uncontrolled behavior and performance degradation. To this end, we propose SCAS, a simple yet effective approach that unifies OOD state correction and OOD action suppression in offline RL. Technically, SCAS achieves value-aware OOD state correction, capable of correcting the agent from OOD states to high-value in-distribution states. Theoretical and empirical results show that SCAS also exhibits the effect of suppressing OOD actions. On standard offline RL benchmarks, SCAS achieves excellent performance without additional hyperparameter tuning. Moreover, benefiting from its OOD state correction feature, SCAS demonstrates enhanced robustness against environmental perturbations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19400
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
Mao, Yixiu
Wang, Qi
Chen, Chen
Qu, Yun
Ji, Xiangyang
Machine Learning
Artificial Intelligence
In offline reinforcement learning (RL), addressing the out-of-distribution (OOD) action issue has been a focus, but we argue that there exists an OOD state issue that also impairs performance yet has been underexplored. Such an issue describes the scenario when the agent encounters states out of the offline dataset during the test phase, leading to uncontrolled behavior and performance degradation. To this end, we propose SCAS, a simple yet effective approach that unifies OOD state correction and OOD action suppression in offline RL. Technically, SCAS achieves value-aware OOD state correction, capable of correcting the agent from OOD states to high-value in-distribution states. Theoretical and empirical results show that SCAS also exhibits the effect of suppressing OOD actions. On standard offline RL benchmarks, SCAS achieves excellent performance without additional hyperparameter tuning. Moreover, benefiting from its OOD state correction feature, SCAS demonstrates enhanced robustness against environmental perturbations.
title Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.19400