Conditioning Matters: Training Diffusion Policies is Faster Than You Think

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Zibin, Liu, Yicheng, Li, Yinchuan, Zhao, Hang, Hao, Jianye
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913842122981376
author Dong, Zibin
Liu, Yicheng
Li, Yinchuan
Zhao, Hang
Hao, Jianye
author_facet Dong, Zibin
Liu, Yicheng
Li, Yinchuan
Zhao, Hang
Hao, Jianye
contents Diffusion policies have emerged as a mainstream paradigm for building vision-language-action (VLA) models. Although they demonstrate strong robot control capabilities, their training efficiency remains suboptimal. In this work, we identify a fundamental challenge in conditional diffusion policy training: when generative conditions are hard to distinguish, the training objective degenerates into modeling the marginal action distribution, a phenomenon we term loss collapse. To overcome this, we propose Cocos, a simple yet general solution that modifies the source distribution in the conditional flow matching to be condition-dependent. By anchoring the source distribution around semantics extracted from condition inputs, Cocos encourages stronger condition integration and prevents the loss collapse. We provide theoretical justification and extensive empirical results across simulation and real-world benchmarks. Our method achieves faster convergence and higher success rates than existing approaches, matching the performance of large-scale pre-trained VLAs using significantly fewer gradient steps and parameters. Cocos is lightweight, easy to implement, and compatible with diverse policy architectures, offering a general-purpose improvement to diffusion policy training.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11123
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Conditioning Matters: Training Diffusion Policies is Faster Than You Think
Dong, Zibin
Liu, Yicheng
Li, Yinchuan
Zhao, Hang
Hao, Jianye
Robotics
Artificial Intelligence
Diffusion policies have emerged as a mainstream paradigm for building vision-language-action (VLA) models. Although they demonstrate strong robot control capabilities, their training efficiency remains suboptimal. In this work, we identify a fundamental challenge in conditional diffusion policy training: when generative conditions are hard to distinguish, the training objective degenerates into modeling the marginal action distribution, a phenomenon we term loss collapse. To overcome this, we propose Cocos, a simple yet general solution that modifies the source distribution in the conditional flow matching to be condition-dependent. By anchoring the source distribution around semantics extracted from condition inputs, Cocos encourages stronger condition integration and prevents the loss collapse. We provide theoretical justification and extensive empirical results across simulation and real-world benchmarks. Our method achieves faster convergence and higher success rates than existing approaches, matching the performance of large-scale pre-trained VLAs using significantly fewer gradient steps and parameters. Cocos is lightweight, easy to implement, and compatible with diverse policy architectures, offering a general-purpose improvement to diffusion policy training.
title Conditioning Matters: Training Diffusion Policies is Faster Than You Think
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2505.11123