Real-Time Operator Takeover for Visuomotor Diffusion Policy Training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Moletta, Marco, Welle, Michael C., Ingelhag, Nils, Munkeby, Jesper, Kragic, Danica
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911556833378304
author Moletta, Marco
Welle, Michael C.
Ingelhag, Nils
Munkeby, Jesper
Kragic, Danica
author_facet Moletta, Marco
Welle, Michael C.
Ingelhag, Nils
Munkeby, Jesper
Kragic, Danica
contents We present a Real-Time Operator Takeover (RTOT) paradigm that enables operators to seamlessly take control of a live visuomotor diffusion policy, guiding the system back to desirable states or providing targeted corrective demonstrations. Within this framework, the operator can intervene to correct the robot's motion, after which control is smoothly returned to the policy until further intervention is needed. We evaluate the takeover framework on three tasks spanning rigid, deformable, and granular objects, and show that incorporating targeted takeover demonstrations significantly improves policy performance compared with training on an equivalent number of initial demonstrations alone. Additionally, we provide an in-depth analysis of the Mahalanobis distance as a signal for automatically identifying undesirable or out-of-distribution states during execution. Supporting materials, including videos of the initial and takeover demonstrations and all experiments, are available on the project website: https://operator-takeover.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2502_02308
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Real-Time Operator Takeover for Visuomotor Diffusion Policy Training
Moletta, Marco
Welle, Michael C.
Ingelhag, Nils
Munkeby, Jesper
Kragic, Danica
Robotics
Machine Learning
We present a Real-Time Operator Takeover (RTOT) paradigm that enables operators to seamlessly take control of a live visuomotor diffusion policy, guiding the system back to desirable states or providing targeted corrective demonstrations. Within this framework, the operator can intervene to correct the robot's motion, after which control is smoothly returned to the policy until further intervention is needed. We evaluate the takeover framework on three tasks spanning rigid, deformable, and granular objects, and show that incorporating targeted takeover demonstrations significantly improves policy performance compared with training on an equivalent number of initial demonstrations alone. Additionally, we provide an in-depth analysis of the Mahalanobis distance as a signal for automatically identifying undesirable or out-of-distribution states during execution. Supporting materials, including videos of the initial and takeover demonstrations and all experiments, are available on the project website: https://operator-takeover.github.io/
title Real-Time Operator Takeover for Visuomotor Diffusion Policy Training
topic Robotics
Machine Learning
url https://arxiv.org/abs/2502.02308