Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Xintong, He, Yutong, Tajwar, Fahim, Salakhutdinov, Ruslan, Kolter, J. Zico, Schneider, Jeff
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912883531579392
author Duan, Xintong
He, Yutong
Tajwar, Fahim
Salakhutdinov, Ruslan
Kolter, J. Zico
Schneider, Jeff
author_facet Duan, Xintong
He, Yutong
Tajwar, Fahim
Salakhutdinov, Ruslan
Kolter, J. Zico
Schneider, Jeff
contents Although diffusion models have achieved strong results in decision-making tasks, their slow inference speed remains a key limitation. While consistency models offer a potential solution, existing applications to decision-making either struggle with suboptimal demonstrations under behavior cloning or rely on complex concurrent training of multiple networks under the actor-critic framework. In this work, we propose a novel approach to consistency distillation for offline reinforcement learning that directly incorporates reward optimization into the distillation process. Our method achieves single-step sampling while generating higher-reward action trajectories through decoupled training and noise-free reward signals. Empirical evaluations on the Gym MuJoCo, FrankaKitchen, and long horizon planning benchmarks demonstrate that our approach can achieve a 9.7% improvement over previous state-of-the-art while offering up to 142x speedup over diffusion counterparts in inference time.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
Duan, Xintong
He, Yutong
Tajwar, Fahim
Salakhutdinov, Ruslan
Kolter, J. Zico
Schneider, Jeff
Machine Learning
Artificial Intelligence
Although diffusion models have achieved strong results in decision-making tasks, their slow inference speed remains a key limitation. While consistency models offer a potential solution, existing applications to decision-making either struggle with suboptimal demonstrations under behavior cloning or rely on complex concurrent training of multiple networks under the actor-critic framework. In this work, we propose a novel approach to consistency distillation for offline reinforcement learning that directly incorporates reward optimization into the distillation process. Our method achieves single-step sampling while generating higher-reward action trajectories through decoupled training and noise-free reward signals. Empirical evaluations on the Gym MuJoCo, FrankaKitchen, and long horizon planning benchmarks demonstrate that our approach can achieve a 9.7% improvement over previous state-of-the-art while offering up to 142x speedup over diffusion counterparts in inference time.
title Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.07822