Fast Visuomotor Policy for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Jingkai, Yang, Tong, Chen, Xueyao, Liu, Chenhuan, Zhang, Wenqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911209374089216
author Jia, Jingkai
Yang, Tong
Chen, Xueyao
Liu, Chenhuan
Zhang, Wenqiang
author_facet Jia, Jingkai
Yang, Tong
Chen, Xueyao
Liu, Chenhuan
Zhang, Wenqiang
contents We present a fast and effective policy framework for robotic manipulation, named Energy Policy, designed for high-frequency robotic tasks and resource-constrained systems. Unlike existing robotic policies, Energy Policy natively predicts multimodal actions in a single forward pass, enabling high-precision manipulation at high speed. The framework is built upon two core components. First, we adopt the energy score as the learning objective to facilitate multimodal action modeling. Second, we introduce an energy MLP to implement the proposed objective while keeping the architecture simple and efficient. We conduct comprehensive experiments in both simulated environments and real-world robotic tasks to evaluate the effectiveness of Energy Policy. The results show that Energy Policy matches or surpasses the performance of state-of-the-art manipulation methods while significantly reducing computational overhead. Notably, on the MimicGen benchmark, Energy Policy achieves superior performance with at a faster inference compared to existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fast Visuomotor Policy for Robotic Manipulation
Jia, Jingkai
Yang, Tong
Chen, Xueyao
Liu, Chenhuan
Zhang, Wenqiang
Robotics
Computer Vision and Pattern Recognition
We present a fast and effective policy framework for robotic manipulation, named Energy Policy, designed for high-frequency robotic tasks and resource-constrained systems. Unlike existing robotic policies, Energy Policy natively predicts multimodal actions in a single forward pass, enabling high-precision manipulation at high speed. The framework is built upon two core components. First, we adopt the energy score as the learning objective to facilitate multimodal action modeling. Second, we introduce an energy MLP to implement the proposed objective while keeping the architecture simple and efficient. We conduct comprehensive experiments in both simulated environments and real-world robotic tasks to evaluate the effectiveness of Energy Policy. The results show that Energy Policy matches or surpasses the performance of state-of-the-art manipulation methods while significantly reducing computational overhead. Notably, on the MimicGen benchmark, Energy Policy achieves superior performance with at a faster inference compared to existing approaches.
title Fast Visuomotor Policy for Robotic Manipulation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.12483