Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Wenbo, Hu, Tianrun, Zhang, Hanbo, Qiao, Yanyuan, Qin, Yuchu, Li, Yang, Liu, Jiajun, Kong, Tao, Liu, Lingqiao, Ma, Xiao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914234630144000
author Zhang, Wenbo
Hu, Tianrun
Zhang, Hanbo
Qiao, Yanyuan
Qin, Yuchu
Li, Yang
Liu, Jiajun
Kong, Tao
Liu, Lingqiao
Ma, Xiao
author_facet Zhang, Wenbo
Hu, Tianrun
Zhang, Hanbo
Qiao, Yanyuan
Qin, Yuchu
Li, Yang
Liu, Jiajun
Kong, Tao
Liu, Lingqiao
Ma, Xiao
contents We present Chain-of-Action (CoA), a novel visuo-motor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-level Chain-of-Thought (CoT) process. This process is unified within a single autoregressive structure: (1) the first token corresponds to a stable keyframe action that encodes the task-specific goals; and (2) subsequent action tokens are generated autoregressively, conditioned on the initial keyframe and previously predicted actions. This backward action reasoning enforces a global-to-local structure, allowing each local action to be tightly constrained by the final goal. To further realize the action reasoning structure, CoA incorporates four complementary designs: continuous action token representation; dynamic stopping for variable-length trajectory generation; reverse temporal ensemble; and multi-token prediction to balance action chunk modeling with global structure. As a result, CoA gives strong spatial generalization capabilities while preserving the flexibility and simplicity of a visuo-motor policy. Empirically, we observe CoA achieves the state-of-the-art performance across 60 RLBench tasks and 8 real-world manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
Zhang, Wenbo
Hu, Tianrun
Zhang, Hanbo
Qiao, Yanyuan
Qin, Yuchu
Li, Yang
Liu, Jiajun
Kong, Tao
Liu, Lingqiao
Ma, Xiao
Robotics
Computer Vision and Pattern Recognition
Machine Learning
We present Chain-of-Action (CoA), a novel visuo-motor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-level Chain-of-Thought (CoT) process. This process is unified within a single autoregressive structure: (1) the first token corresponds to a stable keyframe action that encodes the task-specific goals; and (2) subsequent action tokens are generated autoregressively, conditioned on the initial keyframe and previously predicted actions. This backward action reasoning enforces a global-to-local structure, allowing each local action to be tightly constrained by the final goal. To further realize the action reasoning structure, CoA incorporates four complementary designs: continuous action token representation; dynamic stopping for variable-length trajectory generation; reverse temporal ensemble; and multi-token prediction to balance action chunk modeling with global structure. As a result, CoA gives strong spatial generalization capabilities while preserving the flexibility and simplicity of a visuo-motor policy. Empirically, we observe CoA achieves the state-of-the-art performance across 60 RLBench tasks and 8 real-world manipulation tasks.
title Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2506.09990