Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916262569836544 |
|---|---|
| author | Ackermann, Johannes Osa, Takayuki Sugiyama, Masashi |
| author_facet | Ackermann, Johannes Osa, Takayuki Sugiyama, Masashi |
| contents | Current Reinforcement Learning (RL) is often limited by the large amount of data needed to learn a successful policy. Offline RL aims to solve this issue by using transitions collected by a different behavior policy. We address a novel Offline RL problem setting in which, while collecting the dataset, the transition and reward functions gradually change between episodes but stay constant within each episode. We propose a method based on Contrastive Predictive Coding that identifies this non-stationarity in the offline dataset, accounts for it when training a policy, and predicts it during evaluation. We analyze our proposed method and show that it performs well in simple continuous control tasks and challenging, high-dimensional locomotion tasks. We show that our method often achieves the oracle performance and performs better than baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_14114 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Offline Reinforcement Learning from Datasets with Structured Non-Stationarity Ackermann, Johannes Osa, Takayuki Sugiyama, Masashi Machine Learning Artificial Intelligence Current Reinforcement Learning (RL) is often limited by the large amount of data needed to learn a successful policy. Offline RL aims to solve this issue by using transitions collected by a different behavior policy. We address a novel Offline RL problem setting in which, while collecting the dataset, the transition and reward functions gradually change between episodes but stay constant within each episode. We propose a method based on Contrastive Predictive Coding that identifies this non-stationarity in the offline dataset, accounts for it when training a policy, and predicts it during evaluation. We analyze our proposed method and show that it performs well in simple continuous control tasks and challenging, high-dimensional locomotion tasks. We show that our method often achieves the oracle performance and performs better than baselines. |
| title | Offline Reinforcement Learning from Datasets with Structured Non-Stationarity |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2405.14114 |