Two-Steps Diffusion Policy for Robotic Manipulation via Genetic Denoising

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Clemente, Mateo, Brunswic, Leo, Yang, Rui Heng, Zhao, Xuan, Khalil, Yasser, Lei, Haoyu, Rasouli, Amir, Li, Yinchuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908611406462976
author Clemente, Mateo
Brunswic, Leo
Yang, Rui Heng
Zhao, Xuan
Khalil, Yasser
Lei, Haoyu
Rasouli, Amir
Li, Yinchuan
author_facet Clemente, Mateo
Brunswic, Leo
Yang, Rui Heng
Zhao, Xuan
Khalil, Yasser
Lei, Haoyu
Rasouli, Amir
Li, Yinchuan
contents Diffusion models, such as diffusion policy, have achieved state-of-the-art results in robotic manipulation by imitating expert demonstrations. While diffusion models were originally developed for vision tasks like image and video generation, many of their inference strategies have been directly transferred to control domains without adaptation. In this work, we show that by tailoring the denoising process to the specific characteristics of embodied AI tasks -- particularly structured, low-dimensional nature of action distributions -- diffusion policies can operate effectively with as few as 5 neural function evaluations (NFE). Building on this insight, we propose a population-based sampling strategy, genetic denoising, which enhances both performance and stability by selecting denoising trajectories with low out-of-distribution risk. Our method solves challenging tasks with only 2 NFE while improving or matching performance. We evaluate our approach across 14 robotic manipulation tasks from D4RL and Robomimic, spanning multiple action horizons and inference budgets. In over 2 million evaluations, our method consistently outperforms standard diffusion-based policies, achieving up to 20\% performance gains with significantly fewer inference steps.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Two-Steps Diffusion Policy for Robotic Manipulation via Genetic Denoising
Clemente, Mateo
Brunswic, Leo
Yang, Rui Heng
Zhao, Xuan
Khalil, Yasser
Lei, Haoyu
Rasouli, Amir
Li, Yinchuan
Robotics
Artificial Intelligence
68T40, 93C85, 68T07, 68U35
Diffusion models, such as diffusion policy, have achieved state-of-the-art results in robotic manipulation by imitating expert demonstrations. While diffusion models were originally developed for vision tasks like image and video generation, many of their inference strategies have been directly transferred to control domains without adaptation. In this work, we show that by tailoring the denoising process to the specific characteristics of embodied AI tasks -- particularly structured, low-dimensional nature of action distributions -- diffusion policies can operate effectively with as few as 5 neural function evaluations (NFE). Building on this insight, we propose a population-based sampling strategy, genetic denoising, which enhances both performance and stability by selecting denoising trajectories with low out-of-distribution risk. Our method solves challenging tasks with only 2 NFE while improving or matching performance. We evaluate our approach across 14 robotic manipulation tasks from D4RL and Robomimic, spanning multiple action horizons and inference budgets. In over 2 million evaluations, our method consistently outperforms standard diffusion-based policies, achieving up to 20\% performance gains with significantly fewer inference steps.
title Two-Steps Diffusion Policy for Robotic Manipulation via Genetic Denoising
topic Robotics
Artificial Intelligence
68T40, 93C85, 68T07, 68U35
url https://arxiv.org/abs/2510.21991