VLA Model-Expert Collaboration for Bi-directional Manipulation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Tian-Yu, Jin, Ao-Qun, Zhou, Xiao-Hu, Gui, Mei-Jiang, Xie, Xiao-Liang, Liu, Shi-Qi, Wang, Shuang-Yi, Duang, Sheng-Bin, Wang, Si-Cheng, Lei, Zheng, Hou, Zeng-Guang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912262164316160
author Xiang, Tian-Yu
Jin, Ao-Qun
Zhou, Xiao-Hu
Gui, Mei-Jiang
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Duang, Sheng-Bin
Wang, Si-Cheng
Lei, Zheng
Hou, Zeng-Guang
author_facet Xiang, Tian-Yu
Jin, Ao-Qun
Zhou, Xiao-Hu
Gui, Mei-Jiang
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Duang, Sheng-Bin
Wang, Si-Cheng
Lei, Zheng
Hou, Zeng-Guang
contents The emergence of vision-language-action (VLA) models has given rise to foundation models for robot manipulation. Although these models have achieved significant improvements, their generalization in multi-task manipulation remains limited. This study proposes a VLA model-expert collaboration framework that leverages a limited number of expert actions to enhance VLA model performance. This approach reduces expert workload relative to manual operation while simultaneously improving the reliability and generalization of VLA models. Furthermore, manipulation data collected during collaboration can further refine the VLA model, while human participants concurrently enhance their skills. This bi-directional learning loop boosts the overall performance of the collaboration system. Experimental results across various VLA models demonstrate the effectiveness of the proposed system in collaborative manipulation and learning, as evidenced by improved success rates across tasks. Additionally, validation using a brain-computer interface (BCI) indicates that the collaboration system enhances the efficiency of low-speed action systems by involving VLA model during manipulation. These promising results pave the way for advancing human-robot interaction in the era of foundation models for robotics. (Project website: https://aoqunjin.github.io/Expert-VLA/)
format Preprint
id arxiv_https___arxiv_org_abs_2503_04163
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLA Model-Expert Collaboration for Bi-directional Manipulation Learning
Xiang, Tian-Yu
Jin, Ao-Qun
Zhou, Xiao-Hu
Gui, Mei-Jiang
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Duang, Sheng-Bin
Wang, Si-Cheng
Lei, Zheng
Hou, Zeng-Guang
Robotics
The emergence of vision-language-action (VLA) models has given rise to foundation models for robot manipulation. Although these models have achieved significant improvements, their generalization in multi-task manipulation remains limited. This study proposes a VLA model-expert collaboration framework that leverages a limited number of expert actions to enhance VLA model performance. This approach reduces expert workload relative to manual operation while simultaneously improving the reliability and generalization of VLA models. Furthermore, manipulation data collected during collaboration can further refine the VLA model, while human participants concurrently enhance their skills. This bi-directional learning loop boosts the overall performance of the collaboration system. Experimental results across various VLA models demonstrate the effectiveness of the proposed system in collaborative manipulation and learning, as evidenced by improved success rates across tasks. Additionally, validation using a brain-computer interface (BCI) indicates that the collaboration system enhances the efficiency of low-speed action systems by involving VLA model during manipulation. These promising results pave the way for advancing human-robot interaction in the era of foundation models for robotics. (Project website: https://aoqunjin.github.io/Expert-VLA/)
title VLA Model-Expert Collaboration for Bi-directional Manipulation Learning
topic Robotics
url https://arxiv.org/abs/2503.04163