VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Chongkai, Liu, Zixuan, Chi, Zhenghao, Huang, Junshan, Fei, Xin, Hou, Yiwen, Zhang, Yuxuan, Lin, Yudi, Fang, Zhirui, Jiang, Zeyu, Shao, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908415735889920
author Gao, Chongkai
Liu, Zixuan
Chi, Zhenghao
Huang, Junshan
Fei, Xin
Hou, Yiwen
Zhang, Yuxuan
Lin, Yudi
Fang, Zhirui
Jiang, Zeyu
Shao, Lin
author_facet Gao, Chongkai
Liu, Zixuan
Chi, Zhenghao
Huang, Junshan
Fei, Xin
Hou, Yiwen
Zhang, Yuxuan
Lin, Yudi
Fang, Zhirui
Jiang, Zeyu
Shao, Lin
contents Recent studies on Vision-Language-Action (VLA) models have shifted from the end-to-end action-generation paradigm toward a pipeline involving task planning followed by action generation, demonstrating improved performance on various complex, long-horizon manipulation tasks. However, existing approaches vary significantly in terms of network architectures, planning paradigms, representations, and training data sources, making it challenging for researchers to identify the precise sources of performance gains and components to be further improved. To systematically investigate the impacts of different planning paradigms and representations isolating from network architectures and training data, in this paper, we introduce VLA-OS, a unified VLA architecture series capable of various task planning paradigms, and design a comprehensive suite of controlled experiments across diverse object categories (rigid and deformable), visual modalities (2D and 3D), environments (simulation and real-world), and end-effectors (grippers and dexterous hands). Our results demonstrate that: 1) visually grounded planning representations are generally better than language planning representations; 2) the Hierarchical-VLA paradigm generally achieves superior or comparable performance than other paradigms on task performance, pretraining, generalization ability, scalability, and continual learning ability, albeit at the cost of slower training and inference speeds.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17561
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
Gao, Chongkai
Liu, Zixuan
Chi, Zhenghao
Huang, Junshan
Fei, Xin
Hou, Yiwen
Zhang, Yuxuan
Lin, Yudi
Fang, Zhirui
Jiang, Zeyu
Shao, Lin
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Recent studies on Vision-Language-Action (VLA) models have shifted from the end-to-end action-generation paradigm toward a pipeline involving task planning followed by action generation, demonstrating improved performance on various complex, long-horizon manipulation tasks. However, existing approaches vary significantly in terms of network architectures, planning paradigms, representations, and training data sources, making it challenging for researchers to identify the precise sources of performance gains and components to be further improved. To systematically investigate the impacts of different planning paradigms and representations isolating from network architectures and training data, in this paper, we introduce VLA-OS, a unified VLA architecture series capable of various task planning paradigms, and design a comprehensive suite of controlled experiments across diverse object categories (rigid and deformable), visual modalities (2D and 3D), environments (simulation and real-world), and end-effectors (grippers and dexterous hands). Our results demonstrate that: 1) visually grounded planning representations are generally better than language planning representations; 2) the Hierarchical-VLA paradigm generally achieves superior or comparable performance than other paradigms on task performance, pretraining, generalization ability, scalability, and continual learning ability, albeit at the cost of slower training and inference speeds.
title VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2506.17561