Mini Diffuser: Fast Multi-task Diffusion Policy Training Using Two-level Mini-batches

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hu, Yutong, Song, Pinhao, Wen, Kehan, Detry, Renaud
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915326674862080
author Hu, Yutong
Song, Pinhao
Wen, Kehan
Detry, Renaud
author_facet Hu, Yutong
Song, Pinhao
Wen, Kehan
Detry, Renaud
contents We present a method that reduces, by an order of magnitude, the time and memory needed to train multi-task vision-language robotic diffusion policies. This improvement arises from a previously underexplored distinction between action diffusion and the image diffusion techniques that inspired it: In image generation, the target is high-dimensional. By contrast, in action generation, the dimensionality of the target is comparatively small, and only the image condition is high-dimensional. Our approach, \emph{Mini Diffuser}, exploits this asymmetry by introducing \emph{two-level minibatching}, which pairs multiple noised action samples with each vision-language condition, instead of the conventional one-to-one sampling strategy. To support this batching scheme, we introduce architectural adaptations to the diffusion transformer that prevent information leakage across samples while maintaining full conditioning access. In RLBench simulations, Mini-Diffuser achieves 95\% of the performance of state-of-the-art multi-task diffusion policies, while using only 5\% of the training time and 7\% of the memory. Real-world experiments further validate that Mini-Diffuser preserves the key strengths of diffusion-based policies, including the ability to model multimodal action distributions and produce behavior conditioned on diverse perceptual inputs. Code available at mini-diffuse-actor.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2505_09430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mini Diffuser: Fast Multi-task Diffusion Policy Training Using Two-level Mini-batches
Hu, Yutong
Song, Pinhao
Wen, Kehan
Detry, Renaud
Robotics
Machine Learning
We present a method that reduces, by an order of magnitude, the time and memory needed to train multi-task vision-language robotic diffusion policies. This improvement arises from a previously underexplored distinction between action diffusion and the image diffusion techniques that inspired it: In image generation, the target is high-dimensional. By contrast, in action generation, the dimensionality of the target is comparatively small, and only the image condition is high-dimensional. Our approach, \emph{Mini Diffuser}, exploits this asymmetry by introducing \emph{two-level minibatching}, which pairs multiple noised action samples with each vision-language condition, instead of the conventional one-to-one sampling strategy. To support this batching scheme, we introduce architectural adaptations to the diffusion transformer that prevent information leakage across samples while maintaining full conditioning access. In RLBench simulations, Mini-Diffuser achieves 95\% of the performance of state-of-the-art multi-task diffusion policies, while using only 5\% of the training time and 7\% of the memory. Real-world experiments further validate that Mini-Diffuser preserves the key strengths of diffusion-based policies, including the ability to model multimodal action distributions and produce behavior conditioned on diverse perceptual inputs. Code available at mini-diffuse-actor.github.io
title Mini Diffuser: Fast Multi-task Diffusion Policy Training Using Two-level Mini-batches
topic Robotics
Machine Learning
url https://arxiv.org/abs/2505.09430