Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915859876806656 |
|---|---|
| author | Kim, Jeonghye Lee, Suyoung Kim, Woojun Sung, Youngchul |
| author_facet | Kim, Jeonghye Lee, Suyoung Kim, Woojun Sung, Youngchul |
| contents | Offline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce $Q$-Aided Conditional Supervised Learning (QCS), which effectively combines the stability of RCSL with the stitching capability of $Q$-functions. By analyzing $Q$-function over-generalization, which impairs stable stitching, QCS adaptively integrates $Q$-aid into RCSL's loss function based on trajectory return. Empirical results show that QCS significantly outperforms RCSL and value-based methods, consistently achieving or exceeding the maximum trajectory returns across diverse offline RL benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_02017 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning Kim, Jeonghye Lee, Suyoung Kim, Woojun Sung, Youngchul Machine Learning Offline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce $Q$-Aided Conditional Supervised Learning (QCS), which effectively combines the stability of RCSL with the stitching capability of $Q$-functions. By analyzing $Q$-function over-generalization, which impairs stable stitching, QCS adaptively integrates $Q$-aid into RCSL's loss function based on trajectory return. Empirical results show that QCS significantly outperforms RCSL and value-based methods, consistently achieving or exceeding the maximum trajectory returns across diverse offline RL benchmarks. |
| title | Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2402.02017 |