Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jinouwen, Xue, Rongkun, Niu, Yazhe, Chen, Yun, Yang, Jing, Li, Hongsheng, Liu, Yu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909411766697984
author Zhang, Jinouwen
Xue, Rongkun
Niu, Yazhe
Chen, Yun
Yang, Jing
Li, Hongsheng
Liu, Yu
author_facet Zhang, Jinouwen
Xue, Rongkun
Niu, Yazhe
Chen, Yun
Yang, Jing
Li, Hongsheng
Liu, Yu
contents Generative models, particularly diffusion models, have achieved remarkable success in density estimation for multimodal data, drawing significant interest from the reinforcement learning (RL) community, especially in policy modeling in continuous action spaces. However, existing works exhibit significant variations in training schemes and RL optimization objectives, and some methods are only applicable to diffusion models. In this study, we compare and analyze various generative policy training and deployment techniques, identifying and validating effective designs for generative policy algorithms. Specifically, we revisit existing training objectives and classify them into two categories, each linked to a simpler approach. The first approach, Generative Model Policy Optimization (GMPO), employs a native advantage-weighted regression formulation as the training objective, which is significantly simpler than previous methods. The second approach, Generative Model Policy Gradient (GMPG), offers a numerically stable implementation of the native policy gradient method. We introduce a standardized experimental framework named GenerativeRL. Our experiments demonstrate that the proposed methods achieve state-of-the-art performance on various offline-RL datasets, offering a unified and practical guideline for training and deploying generative policies.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01245
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
Zhang, Jinouwen
Xue, Rongkun
Niu, Yazhe
Chen, Yun
Yang, Jing
Li, Hongsheng
Liu, Yu
Machine Learning
Artificial Intelligence
Generative models, particularly diffusion models, have achieved remarkable success in density estimation for multimodal data, drawing significant interest from the reinforcement learning (RL) community, especially in policy modeling in continuous action spaces. However, existing works exhibit significant variations in training schemes and RL optimization objectives, and some methods are only applicable to diffusion models. In this study, we compare and analyze various generative policy training and deployment techniques, identifying and validating effective designs for generative policy algorithms. Specifically, we revisit existing training objectives and classify them into two categories, each linked to a simpler approach. The first approach, Generative Model Policy Optimization (GMPO), employs a native advantage-weighted regression formulation as the training objective, which is significantly simpler than previous methods. The second approach, Generative Model Policy Gradient (GMPG), offers a numerically stable implementation of the native policy gradient method. We introduce a standardized experimental framework named GenerativeRL. Our experiments demonstrate that the proposed methods achieve state-of-the-art performance on various offline-RL datasets, offering a unified and practical guideline for training and deploying generative policies.
title Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2412.01245