ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xia, Wenlong, Zhang, Jinhao, Zhang, Ce, Wang, Yaojia, Li, Huizhe, Gong, Youmin, Mei, Jie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908858076626944
author Xia, Wenlong
Zhang, Jinhao
Zhang, Ce
Wang, Yaojia
Li, Huizhe
Gong, Youmin
Mei, Jie
author_facet Xia, Wenlong
Zhang, Jinhao
Zhang, Ce
Wang, Yaojia
Li, Huizhe
Gong, Youmin
Mei, Jie
contents Vision-based imitation learning has enabled impressive robotic manipulation skills, but its reliance on object appearance while ignoring the underlying 3D scene structure leads to low training efficiency and poor generalization. To address these challenges, we introduce \emph{Implicit Scene Supervision (ISS) Policy}, a 3D visuomotor DiT-based diffusion policy that predicts sequences of continuous actions from point cloud observations. We extend DiT with a novel implicit scene supervision module that encourages the model to produce outputs consistent with the scene's geometric evolution, thereby improving the performance and robustness of the policy. Notably, ISS Policy achieves state-of-the-art performance on both single-arm manipulation tasks (MetaWorld) and dexterous hand manipulation (Adroit). In real-world experiments, it also demonstrates strong generalization and robustness. Additional ablation studies show that our method scales effectively with both data and parameters. Code and videos will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15020
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision
Xia, Wenlong
Zhang, Jinhao
Zhang, Ce
Wang, Yaojia
Li, Huizhe
Gong, Youmin
Mei, Jie
Robotics
Vision-based imitation learning has enabled impressive robotic manipulation skills, but its reliance on object appearance while ignoring the underlying 3D scene structure leads to low training efficiency and poor generalization. To address these challenges, we introduce \emph{Implicit Scene Supervision (ISS) Policy}, a 3D visuomotor DiT-based diffusion policy that predicts sequences of continuous actions from point cloud observations. We extend DiT with a novel implicit scene supervision module that encourages the model to produce outputs consistent with the scene's geometric evolution, thereby improving the performance and robustness of the policy. Notably, ISS Policy achieves state-of-the-art performance on both single-arm manipulation tasks (MetaWorld) and dexterous hand manipulation (Adroit). In real-world experiments, it also demonstrates strong generalization and robustness. Additional ablation studies show that our method scales effectively with both data and parameters. Code and videos will be released.
title ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision
topic Robotics
url https://arxiv.org/abs/2512.15020