Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Mingyang, Ding, Pengxiang, Zhang, Weinan, Wang, Donglin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912239718498304
author Sun, Mingyang
Ding, Pengxiang
Zhang, Weinan
Wang, Donglin
author_facet Sun, Mingyang
Ding, Pengxiang
Zhang, Weinan
Wang, Donglin
contents Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and long-term planning. However, they face challenges in robustness when encountering distribution shifts. This paper explores improving diffusion-based imitation learning models through online interactions with the environment. We propose OTPR (Optimal Transport-guided score-based diffusion Policy for Reinforcement learning fine-tuning), a novel method that integrates diffusion policies with RL using optimal transport theory. OTPR leverages the Q-function as a transport cost and views the policy as an optimal transport map, enabling efficient and stable fine-tuning. Moreover, we introduce masked optimal transport to guide state-action matching using expert keypoints and a compatibility-based resampling strategy to enhance training stability. Experiments on three simulation tasks demonstrate OTPR's superior performance and robustness compared to existing methods, especially in complex and sparse-reward environments. In sum, OTPR provides an effective framework for combining IL and RL, achieving versatile and reliable policy learning. The code will be released at https://github.com/Sunmmyy/OTPR.git.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
Sun, Mingyang
Ding, Pengxiang
Zhang, Weinan
Wang, Donglin
Machine Learning
Artificial Intelligence
Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and long-term planning. However, they face challenges in robustness when encountering distribution shifts. This paper explores improving diffusion-based imitation learning models through online interactions with the environment. We propose OTPR (Optimal Transport-guided score-based diffusion Policy for Reinforcement learning fine-tuning), a novel method that integrates diffusion policies with RL using optimal transport theory. OTPR leverages the Q-function as a transport cost and views the policy as an optimal transport map, enabling efficient and stable fine-tuning. Moreover, we introduce masked optimal transport to guide state-action matching using expert keypoints and a compatibility-based resampling strategy to enhance training stability. Experiments on three simulation tasks demonstrate OTPR's superior performance and robustness compared to existing methods, especially in complex and sparse-reward environments. In sum, OTPR provides an effective framework for combining IL and RL, achieving versatile and reliable policy learning. The code will be released at https://github.com/Sunmmyy/OTPR.git.
title Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.12631