CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Zhefei, Ding, Pengxiang, Lyu, Shangke, Huang, Siteng, Sun, Mingyang, Zhao, Wei, Fan, Zhaoxin, Wang, Donglin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916888890572800
author Gong, Zhefei
Ding, Pengxiang
Lyu, Shangke
Huang, Siteng
Sun, Mingyang
Zhao, Wei
Fan, Zhaoxin
Wang, Donglin
author_facet Gong, Zhefei
Ding, Pengxiang
Lyu, Shangke
Huang, Siteng
Sun, Mingyang
Zhao, Wei
Fan, Zhaoxin
Wang, Donglin
contents In robotic visuomotor policy learning, diffusion-based models have achieved significant success in improving the accuracy of action trajectory generation compared to traditional autoregressive models. However, they suffer from inefficiency due to multiple denoising steps and limited flexibility from complex constraints. In this paper, we introduce Coarse-to-Fine AutoRegressive Policy (CARP), a novel paradigm for visuomotor policy learning that redefines the autoregressive action generation process as a coarse-to-fine, next-scale approach. CARP decouples action generation into two stages: first, an action autoencoder learns multi-scale representations of the entire action sequence; then, a GPT-style transformer refines the sequence prediction through a coarse-to-fine autoregressive process. This straightforward and intuitive approach produces highly accurate and smooth actions, matching or even surpassing the performance of diffusion-based policies while maintaining efficiency on par with autoregressive policies. We conduct extensive evaluations across diverse settings, including single-task and multi-task scenarios on state-based and image-based simulation benchmarks, as well as real-world tasks. CARP achieves competitive success rates, with up to a 10% improvement, and delivers 10x faster inference compared to state-of-the-art policies, establishing a high-performance, efficient, and flexible paradigm for action generation in robotic tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06782
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
Gong, Zhefei
Ding, Pengxiang
Lyu, Shangke
Huang, Siteng
Sun, Mingyang
Zhao, Wei
Fan, Zhaoxin
Wang, Donglin
Robotics
Computer Vision and Pattern Recognition
In robotic visuomotor policy learning, diffusion-based models have achieved significant success in improving the accuracy of action trajectory generation compared to traditional autoregressive models. However, they suffer from inefficiency due to multiple denoising steps and limited flexibility from complex constraints. In this paper, we introduce Coarse-to-Fine AutoRegressive Policy (CARP), a novel paradigm for visuomotor policy learning that redefines the autoregressive action generation process as a coarse-to-fine, next-scale approach. CARP decouples action generation into two stages: first, an action autoencoder learns multi-scale representations of the entire action sequence; then, a GPT-style transformer refines the sequence prediction through a coarse-to-fine autoregressive process. This straightforward and intuitive approach produces highly accurate and smooth actions, matching or even surpassing the performance of diffusion-based policies while maintaining efficiency on par with autoregressive policies. We conduct extensive evaluations across diverse settings, including single-task and multi-task scenarios on state-based and image-based simulation benchmarks, as well as real-world tasks. CARP achieves competitive success rates, with up to a 10% improvement, and delivers 10x faster inference compared to state-of-the-art policies, establishing a high-performance, efficient, and flexible paradigm for action generation in robotic tasks.
title CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.06782