Forward KL Regularized Preference Optimization for Aligning Diffusion Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shan, Zhao, Fan, Chenyou, Qiu, Shuang, Shi, Jiyuan, Bai, Chenjia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929631908593664
author Shan, Zhao
Fan, Chenyou
Qiu, Shuang
Shi, Jiyuan
Bai, Chenjia
author_facet Shan, Zhao
Fan, Chenyou
Qiu, Shuang
Shi, Jiyuan
Bai, Chenjia
contents Diffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning. A central problem for learning diffusion policies is to align the policy output with human intents in various tasks. To achieve this, previous methods conduct return-conditioned policy generation or Reinforcement Learning (RL)-based policy optimization, while they both rely on pre-defined reward functions. In this work, we propose a novel framework, Forward KL regularized Preference optimization for aligning Diffusion policies, to align the diffusion policy with preferences directly. We first train a diffusion policy from the offline dataset without considering the preference, and then align the policy to the preference data via direct preference optimization. During the alignment phase, we formulate direct preference learning in a diffusion policy, where the forward KL regularization is employed in preference optimization to avoid generating out-of-distribution actions. We conduct extensive experiments for MetaWorld manipulation and D4RL tasks. The results show our method exhibits superior alignment with preferences and outperforms previous state-of-the-art algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2409_05622
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Forward KL Regularized Preference Optimization for Aligning Diffusion Policies
Shan, Zhao
Fan, Chenyou
Qiu, Shuang
Shi, Jiyuan
Bai, Chenjia
Machine Learning
Diffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning. A central problem for learning diffusion policies is to align the policy output with human intents in various tasks. To achieve this, previous methods conduct return-conditioned policy generation or Reinforcement Learning (RL)-based policy optimization, while they both rely on pre-defined reward functions. In this work, we propose a novel framework, Forward KL regularized Preference optimization for aligning Diffusion policies, to align the diffusion policy with preferences directly. We first train a diffusion policy from the offline dataset without considering the preference, and then align the policy to the preference data via direct preference optimization. During the alignment phase, we formulate direct preference learning in a diffusion policy, where the forward KL regularization is employed in preference optimization to avoid generating out-of-distribution actions. We conduct extensive experiments for MetaWorld manipulation and D4RL tasks. The results show our method exhibits superior alignment with preferences and outperforms previous state-of-the-art algorithms.
title Forward KL Regularized Preference Optimization for Aligning Diffusion Policies
topic Machine Learning
url https://arxiv.org/abs/2409.05622