Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912883531579392 |
|---|---|
| author | Duan, Xintong He, Yutong Tajwar, Fahim Salakhutdinov, Ruslan Kolter, J. Zico Schneider, Jeff |
| author_facet | Duan, Xintong He, Yutong Tajwar, Fahim Salakhutdinov, Ruslan Kolter, J. Zico Schneider, Jeff |
| contents | Although diffusion models have achieved strong results in decision-making tasks, their slow inference speed remains a key limitation. While consistency models offer a potential solution, existing applications to decision-making either struggle with suboptimal demonstrations under behavior cloning or rely on complex concurrent training of multiple networks under the actor-critic framework. In this work, we propose a novel approach to consistency distillation for offline reinforcement learning that directly incorporates reward optimization into the distillation process. Our method achieves single-step sampling while generating higher-reward action trajectories through decoupled training and noise-free reward signals. Empirical evaluations on the Gym MuJoCo, FrankaKitchen, and long horizon planning benchmarks demonstrate that our approach can achieve a 9.7% improvement over previous state-of-the-art while offering up to 142x speedup over diffusion counterparts in inference time. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_07822 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation Duan, Xintong He, Yutong Tajwar, Fahim Salakhutdinov, Ruslan Kolter, J. Zico Schneider, Jeff Machine Learning Artificial Intelligence Although diffusion models have achieved strong results in decision-making tasks, their slow inference speed remains a key limitation. While consistency models offer a potential solution, existing applications to decision-making either struggle with suboptimal demonstrations under behavior cloning or rely on complex concurrent training of multiple networks under the actor-critic framework. In this work, we propose a novel approach to consistency distillation for offline reinforcement learning that directly incorporates reward optimization into the distillation process. Our method achieves single-step sampling while generating higher-reward action trajectories through decoupled training and noise-free reward signals. Empirical evaluations on the Gym MuJoCo, FrankaKitchen, and long horizon planning benchmarks demonstrate that our approach can achieve a 9.7% improvement over previous state-of-the-art while offering up to 142x speedup over diffusion counterparts in inference time. |
| title | Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2506.07822 |