Saved in:
Bibliographic Details
Main Authors: Wu, Wei, Lu, Fan, Wang, Yunnan, Yang, Shuai, Liu, Shi, Wang, Fangjing, Zhu, Qian, Sun, He, Wang, Yong, Ma, Shuailei, Ren, Yiyu, Zhang, Kejia, Yu, Hui, Zhao, Jingmei, Zhou, Shuai, Qiu, Zhenqi, Xiong, Houlong, Wang, Ziyu, Wang, Zechen, Cheng, Ran, Li, Yong-Lu, Huang, Yongtao, Zhu, Xing, Shen, Yujun, Zheng, Kecheng
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.18692
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910033400299520
author Wu, Wei
Lu, Fan
Wang, Yunnan
Yang, Shuai
Liu, Shi
Wang, Fangjing
Zhu, Qian
Sun, He
Wang, Yong
Ma, Shuailei
Ren, Yiyu
Zhang, Kejia
Yu, Hui
Zhao, Jingmei
Zhou, Shuai
Qiu, Zhenqi
Xiong, Houlong
Wang, Ziyu
Wang, Zechen
Cheng, Ran
Li, Yong-Lu
Huang, Yongtao
Zhu, Xing
Shen, Yujun
Zheng, Kecheng
author_facet Wu, Wei
Lu, Fan
Wang, Yunnan
Yang, Shuai
Liu, Shi
Wang, Fangjing
Zhu, Qian
Sun, He
Wang, Yong
Ma, Shuailei
Ren, Yiyu
Zhang, Kejia
Yu, Hui
Zhao, Jingmei
Zhou, Shuai
Qiu, Zhenqi
Xiong, Houlong
Wang, Ziyu
Wang, Zechen
Cheng, Ran
Li, Yong-Lu
Huang, Yongtao
Zhu, Xing
Shen, Yujun
Zheng, Kecheng
contents Offering great potential in robotic manipulation, a capable Vision-Language-Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensuring cost efficiency (e.g., data and GPU hours required for adaptation). To this end, we develop LingBot-VLA with around 20,000 hours of real-world data from 9 popular dual-arm robot configurations. Through a systematic assessment on 3 robotic platforms, each completing 100 tasks with 130 post-training episodes per task, our model achieves clear superiority over competitors, showcasing its strong performance and broad generalizability. We have also built an efficient codebase, which delivers a throughput of 261 samples per second with an 8-GPU training setup, representing a 1.5~2.8$\times$ (depending on the relied VLM base model) speedup over existing VLA-oriented codebases. The above features ensure that our model is well-suited for real-world deployment. To advance the field of robot learning, we provide open access to the code, base model, and benchmark data, with a focus on enabling more challenging tasks and promoting sound evaluation standards.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18692
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Pragmatic VLA Foundation Model
Wu, Wei
Lu, Fan
Wang, Yunnan
Yang, Shuai
Liu, Shi
Wang, Fangjing
Zhu, Qian
Sun, He
Wang, Yong
Ma, Shuailei
Ren, Yiyu
Zhang, Kejia
Yu, Hui
Zhao, Jingmei
Zhou, Shuai
Qiu, Zhenqi
Xiong, Houlong
Wang, Ziyu
Wang, Zechen
Cheng, Ran
Li, Yong-Lu
Huang, Yongtao
Zhu, Xing
Shen, Yujun
Zheng, Kecheng
Robotics
Computer Vision and Pattern Recognition
Offering great potential in robotic manipulation, a capable Vision-Language-Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensuring cost efficiency (e.g., data and GPU hours required for adaptation). To this end, we develop LingBot-VLA with around 20,000 hours of real-world data from 9 popular dual-arm robot configurations. Through a systematic assessment on 3 robotic platforms, each completing 100 tasks with 130 post-training episodes per task, our model achieves clear superiority over competitors, showcasing its strong performance and broad generalizability. We have also built an efficient codebase, which delivers a throughput of 261 samples per second with an 8-GPU training setup, representing a 1.5~2.8$\times$ (depending on the relied VLM base model) speedup over existing VLA-oriented codebases. The above features ensure that our model is well-suited for real-world deployment. To advance the field of robot learning, we provide open access to the code, base model, and benchmark data, with a focus on enabling more challenging tasks and promoting sound evaluation standards.
title A Pragmatic VLA Foundation Model
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.18692