Enhanced DACER Algorithm with High Diffusion Efficiency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yinuo, Wang, Likun, Tan, Mining, Zou, Wenjun, Song, Xujie, Wang, Wenxuan, Liu, Tong, Zhan, Guojian, Zhu, Tianze, Liu, Shiqi, He, Zeyu, Zhang, Feihong, Duan, Jingliang, Li, Shengbo Eben
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908572570353664
author Wang, Yinuo
Wang, Likun
Tan, Mining
Zou, Wenjun
Song, Xujie
Wang, Wenxuan
Liu, Tong
Zhan, Guojian
Zhu, Tianze
Liu, Shiqi
He, Zeyu
Zhang, Feihong
Duan, Jingliang
Li, Shengbo Eben
author_facet Wang, Yinuo
Wang, Likun
Tan, Mining
Zou, Wenjun
Song, Xujie
Wang, Wenxuan
Liu, Tong
Zhan, Guojian
Zhu, Tianze
Liu, Shiqi
He, Zeyu
Zhang, Feihong
Duan, Jingliang
Li, Shengbo Eben
contents Due to their expressive capacity, diffusion models have shown great promise in offline RL and imitation learning. Diffusion Actor-Critic with Entropy Regulator (DACER) extended this capability to online RL by using the reverse diffusion process as a policy approximator, achieving state-of-the-art performance. However, it still suffers from a core trade-off: more diffusion steps ensure high performance but reduce efficiency, while fewer steps degrade performance. This remains a major bottleneck for deploying diffusion policies in real-time online RL. To mitigate this, we propose DACERv2, which leverages a Q-gradient field objective with respect to action as an auxiliary optimization target to guide the denoising process at each diffusion step, thereby introducing intermediate supervisory signals that enhance the efficiency of single-step diffusion. Additionally, we observe that the independence of the Q-gradient field from the diffusion time step is inconsistent with the characteristics of the diffusion process. To address this issue, a temporal weighting mechanism is introduced, allowing the model to effectively eliminate large-scale noise during the early stages and refine its outputs in the later stages. Experimental results on OpenAI Gym benchmarks and multimodal tasks demonstrate that, compared with classical and diffusion-based online RL algorithms, DACERv2 achieves higher performance in most complex control environments with only five diffusion steps and shows greater multimodality.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23426
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhanced DACER Algorithm with High Diffusion Efficiency
Wang, Yinuo
Wang, Likun
Tan, Mining
Zou, Wenjun
Song, Xujie
Wang, Wenxuan
Liu, Tong
Zhan, Guojian
Zhu, Tianze
Liu, Shiqi
He, Zeyu
Zhang, Feihong
Duan, Jingliang
Li, Shengbo Eben
Machine Learning
Artificial Intelligence
Due to their expressive capacity, diffusion models have shown great promise in offline RL and imitation learning. Diffusion Actor-Critic with Entropy Regulator (DACER) extended this capability to online RL by using the reverse diffusion process as a policy approximator, achieving state-of-the-art performance. However, it still suffers from a core trade-off: more diffusion steps ensure high performance but reduce efficiency, while fewer steps degrade performance. This remains a major bottleneck for deploying diffusion policies in real-time online RL. To mitigate this, we propose DACERv2, which leverages a Q-gradient field objective with respect to action as an auxiliary optimization target to guide the denoising process at each diffusion step, thereby introducing intermediate supervisory signals that enhance the efficiency of single-step diffusion. Additionally, we observe that the independence of the Q-gradient field from the diffusion time step is inconsistent with the characteristics of the diffusion process. To address this issue, a temporal weighting mechanism is introduced, allowing the model to effectively eliminate large-scale noise during the early stages and refine its outputs in the later stages. Experimental results on OpenAI Gym benchmarks and multimodal tasks demonstrate that, compared with classical and diffusion-based online RL algorithms, DACERv2 achieves higher performance in most complex control environments with only five diffusion steps and shows greater multimodality.
title Enhanced DACER Algorithm with High Diffusion Efficiency
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.23426