A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chenxuan, Liu, Jiaming, Wang, Guanqun, Li, Xiaoqi, Chen, Sixiang, Heng, Liang, Xiong, Chuyan, Ge, Jiaxin, Zhang, Renrui, Zhou, Kaichen, Zhang, Shanghang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917960412561408
author Li, Chenxuan
Liu, Jiaming
Wang, Guanqun
Li, Xiaoqi
Chen, Sixiang
Heng, Liang
Xiong, Chuyan
Ge, Jiaxin
Zhang, Renrui
Zhou, Kaichen
Zhang, Shanghang
author_facet Li, Chenxuan
Liu, Jiaming
Wang, Guanqun
Li, Xiaoqi
Chen, Sixiang
Heng, Liang
Xiong, Chuyan
Ge, Jiaxin
Zhang, Renrui
Zhou, Kaichen
Zhang, Shanghang
contents Recently, some studies have integrated Multimodal Large Language Models into robotic manipulation, constructing vision-language-action models (VLAs) to interpret multimodal information and predict SE(3) poses. While VLAs have shown promising progress, they may suffer from failures when faced with novel and complex tasks. To emulate human-like reasoning for more robust manipulation, we propose the self-corrected (SC-)VLA framework, which integrates fast system for directly predicting actions and slow system for reflecting on failed actions within a single VLA policy. For the fast system, we incorporate parameter-efficient fine-tuning to equip the model with pose prediction capabilities while preserving the inherent reasoning abilities of MLLMs. For the slow system, we propose a Chain-of-Thought training strategy for failure correction, designed to mimic human reflection after a manipulation failure. Specifically, our model learns to identify the causes of action failures, adaptively seek expert feedback, reflect on the current failure scenario, and iteratively generate corrective actions, step by step. Furthermore, a continuous policy learning method is designed based on successfully corrected samples, enhancing the fast system's adaptability to the current configuration. We compare SC-VLA with the previous SOTA VLA in both simulation and real-world tasks, demonstrating an efficient correction process and improved manipulation accuracy on both seen and unseen tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17418
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
Li, Chenxuan
Liu, Jiaming
Wang, Guanqun
Li, Xiaoqi
Chen, Sixiang
Heng, Liang
Xiong, Chuyan
Ge, Jiaxin
Zhang, Renrui
Zhou, Kaichen
Zhang, Shanghang
Computer Vision and Pattern Recognition
Recently, some studies have integrated Multimodal Large Language Models into robotic manipulation, constructing vision-language-action models (VLAs) to interpret multimodal information and predict SE(3) poses. While VLAs have shown promising progress, they may suffer from failures when faced with novel and complex tasks. To emulate human-like reasoning for more robust manipulation, we propose the self-corrected (SC-)VLA framework, which integrates fast system for directly predicting actions and slow system for reflecting on failed actions within a single VLA policy. For the fast system, we incorporate parameter-efficient fine-tuning to equip the model with pose prediction capabilities while preserving the inherent reasoning abilities of MLLMs. For the slow system, we propose a Chain-of-Thought training strategy for failure correction, designed to mimic human reflection after a manipulation failure. Specifically, our model learns to identify the causes of action failures, adaptively seek expert feedback, reflect on the current failure scenario, and iteratively generate corrective actions, step by step. Furthermore, a continuous policy learning method is designed based on successfully corrected samples, enhancing the fast system's adaptability to the current configuration. We compare SC-VLA with the previous SOTA VLA in both simulation and real-world tasks, demonstrating an efficient correction process and improved manipulation accuracy on both seen and unseen tasks.
title A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.17418