Saved in:
Bibliographic Details
Main Authors: Shan, Yixiang, Zhu, Zhengbang, Long, Ting, Liang, Qifan, Chang, Yi, Zhang, Weinan, Yin, Liang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.02772
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917694573379584
author Shan, Yixiang
Zhu, Zhengbang
Long, Ting
Liang, Qifan
Chang, Yi
Zhang, Weinan
Yin, Liang
author_facet Shan, Yixiang
Zhu, Zhengbang
Long, Ting
Liang, Qifan
Chang, Yi
Zhang, Weinan
Yin, Liang
contents The performance of offline reinforcement learning (RL) is sensitive to the proportion of high-return trajectories in the offline dataset. However, in many simulation environments and real-world scenarios, there are large ratios of low-return trajectories rather than high-return trajectories, which makes learning an efficient policy challenging. In this paper, we propose a method called Contrastive Diffuser (CDiffuser) to make full use of low-return trajectories and improve the performance of offline RL algorithms. Specifically, CDiffuser groups the states of trajectories in the offline dataset into high-return states and low-return states and treats them as positive and negative samples correspondingly. Then, it designs a contrastive mechanism to pull the trajectory of an agent toward high-return states and push them away from low-return states. Through the contrast mechanism, trajectories with low returns can serve as negative examples for policy learning, guiding the agent to avoid areas associated with low returns and achieve better performance. Experiments on 14 commonly used D4RL benchmarks demonstrate the effectiveness of our proposed method. Our code is publicly available at \url{https://anonymous.4open.science/r/CDiffuser}.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02772
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contrastive Diffuser: Planning Towards High Return States via Contrastive Learning
Shan, Yixiang
Zhu, Zhengbang
Long, Ting
Liang, Qifan
Chang, Yi
Zhang, Weinan
Yin, Liang
Machine Learning
The performance of offline reinforcement learning (RL) is sensitive to the proportion of high-return trajectories in the offline dataset. However, in many simulation environments and real-world scenarios, there are large ratios of low-return trajectories rather than high-return trajectories, which makes learning an efficient policy challenging. In this paper, we propose a method called Contrastive Diffuser (CDiffuser) to make full use of low-return trajectories and improve the performance of offline RL algorithms. Specifically, CDiffuser groups the states of trajectories in the offline dataset into high-return states and low-return states and treats them as positive and negative samples correspondingly. Then, it designs a contrastive mechanism to pull the trajectory of an agent toward high-return states and push them away from low-return states. Through the contrast mechanism, trajectories with low returns can serve as negative examples for policy learning, guiding the agent to avoid areas associated with low returns and achieve better performance. Experiments on 14 commonly used D4RL benchmarks demonstrate the effectiveness of our proposed method. Our code is publicly available at \url{https://anonymous.4open.science/r/CDiffuser}.
title Contrastive Diffuser: Planning Towards High Return States via Contrastive Learning
topic Machine Learning
url https://arxiv.org/abs/2402.02772