Any-step Dynamics Model Improves Future Predictions for Online and Offline Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916261936496640 |
|---|---|
| author | Lin, Haoxin Xu, Yu-Yan Sun, Yihao Zhang, Zhilong Li, Yi-Chen Jia, Chengxing Ye, Junyin Zhang, Jiaji Yu, Yang |
| author_facet | Lin, Haoxin Xu, Yu-Yan Sun, Yihao Zhang, Zhilong Li, Yi-Chen Jia, Chengxing Ye, Junyin Zhang, Jiaji Yu, Yang |
| contents | Model-based methods in reinforcement learning offer a promising approach to enhance data efficiency by facilitating policy exploration within a dynamics model. However, accurately predicting sequential steps in the dynamics model remains a challenge due to the bootstrapping prediction, which attributes the next state to the prediction of the current state. This leads to accumulated errors during model roll-out. In this paper, we propose the Any-step Dynamics Model (ADM) to mitigate the compounding error by reducing bootstrapping prediction to direct prediction. ADM allows for the use of variable-length plans as inputs for predicting future states without frequent bootstrapping. We design two algorithms, ADMPO-ON and ADMPO-OFF, which apply ADM in online and offline model-based frameworks, respectively. In the online setting, ADMPO-ON demonstrates improved sample efficiency compared to previous state-of-the-art methods. In the offline setting, ADMPO-OFF not only demonstrates superior performance compared to recent state-of-the-art offline approaches but also offers better quantification of model uncertainty using only a single ADM. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_17031 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Any-step Dynamics Model Improves Future Predictions for Online and Offline Reinforcement Learning Lin, Haoxin Xu, Yu-Yan Sun, Yihao Zhang, Zhilong Li, Yi-Chen Jia, Chengxing Ye, Junyin Zhang, Jiaji Yu, Yang Machine Learning Model-based methods in reinforcement learning offer a promising approach to enhance data efficiency by facilitating policy exploration within a dynamics model. However, accurately predicting sequential steps in the dynamics model remains a challenge due to the bootstrapping prediction, which attributes the next state to the prediction of the current state. This leads to accumulated errors during model roll-out. In this paper, we propose the Any-step Dynamics Model (ADM) to mitigate the compounding error by reducing bootstrapping prediction to direct prediction. ADM allows for the use of variable-length plans as inputs for predicting future states without frequent bootstrapping. We design two algorithms, ADMPO-ON and ADMPO-OFF, which apply ADM in online and offline model-based frameworks, respectively. In the online setting, ADMPO-ON demonstrates improved sample efficiency compared to previous state-of-the-art methods. In the offline setting, ADMPO-OFF not only demonstrates superior performance compared to recent state-of-the-art offline approaches but also offers better quantification of model uncertainty using only a single ADM. |
| title | Any-step Dynamics Model Improves Future Predictions for Online and Offline Reinforcement Learning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2405.17031 |