Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Hao, Feng, Yicheng, Zhang, Wanpeng, Zheng, Sipeng, Wang, Ye, Yuan, Haoqi, Liu, Jiazheng, Xu, Chaoyi, Jin, Qin, Lu, Zongqing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911067536359424
author Luo, Hao
Feng, Yicheng
Zhang, Wanpeng
Zheng, Sipeng
Wang, Ye
Yuan, Haoqi
Liu, Jiazheng
Xu, Chaoyi
Jin, Qin
Lu, Zongqing
author_facet Luo, Hao
Feng, Yicheng
Zhang, Wanpeng
Zheng, Sipeng
Wang, Ye
Yuan, Haoqi
Liu, Jiazheng
Xu, Chaoyi
Jin, Qin
Lu, Zongqing
contents We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks, primarily due to their reliance on synthetic data with significant sim-to-real gaps or teleoperated demonstrations lacking scale and diversity. To address this data bottleneck, we propose leveraging human hands as a foundation manipulator, capitalizing on the rich dexterity and scalability present in web data. Our approach centers on physical instruction tuning, a novel training paradigm that combines large-scale VLA pretraining from human videos, physical space alignment for 3D reasoning, and post-training adaptation for robotic tasks. Additionally, we introduce a part-level motion tokenization method which achieves millimeter-level reconstruction accuracy to model precise hand trajectories for action learning. To support our proposed paradigm, we further develop a comprehensive data curation pipeline that integrates heterogeneous sources -- including motion capture, VR, and RGB-only videos -- into a large-scale dataset with millions of motion-based instructional instances. We empirically show the excellence of Being-H0 in hand motion generation and instruction following, and it also scales well with model and data sizes. Importantly, we observe the expected gains of Being-H0 in real-world robotic manipulation as physical instruction tuning is applied. More details are available at https://beingbeyond.github.io/Being-H0.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15597
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Luo, Hao
Feng, Yicheng
Zhang, Wanpeng
Zheng, Sipeng
Wang, Ye
Yuan, Haoqi
Liu, Jiazheng
Xu, Chaoyi
Jin, Qin
Lu, Zongqing
Computer Vision and Pattern Recognition
Machine Learning
Robotics
We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks, primarily due to their reliance on synthetic data with significant sim-to-real gaps or teleoperated demonstrations lacking scale and diversity. To address this data bottleneck, we propose leveraging human hands as a foundation manipulator, capitalizing on the rich dexterity and scalability present in web data. Our approach centers on physical instruction tuning, a novel training paradigm that combines large-scale VLA pretraining from human videos, physical space alignment for 3D reasoning, and post-training adaptation for robotic tasks. Additionally, we introduce a part-level motion tokenization method which achieves millimeter-level reconstruction accuracy to model precise hand trajectories for action learning. To support our proposed paradigm, we further develop a comprehensive data curation pipeline that integrates heterogeneous sources -- including motion capture, VR, and RGB-only videos -- into a large-scale dataset with millions of motion-based instructional instances. We empirically show the excellence of Being-H0 in hand motion generation and instruction following, and it also scales well with model and data sizes. Importantly, we observe the expected gains of Being-H0 in real-world robotic manipulation as physical instruction tuning is applied. More details are available at https://beingbeyond.github.io/Being-H0.
title Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
topic Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2507.15597