Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Hao, Feng, Yicheng, Zhang, Wanpeng, Zheng, Sipeng, Wang, Ye, Yuan, Haoqi, Liu, Jiazheng, Xu, Chaoyi, Jin, Qin, Lu, Zongqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
by: Feng, Yicheng, et al.
Published: (2025)
by: Feng, Yicheng, et al.
Published: (2025)
Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Conservative Offline Robot Policy Learning via Posterior-Transition Reweighting
by: Zhang, Wanpeng, et al.
Published: (2026)
by: Zhang, Wanpeng, et al.
Published: (2026)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
by: Xu, Haiweng, et al.
Published: (2026)
by: Xu, Haiweng, et al.
Published: (2026)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
by: Cao, Bin, et al.
Published: (2025)
by: Cao, Bin, et al.
Published: (2025)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)
by: Zheng, Sipeng, et al.
Published: (2024)
Scaling Large Motion Models with Million-Level Human Motions
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
by: Yuan, Haoqi, et al.
Published: (2025)
by: Yuan, Haoqi, et al.
Published: (2025)
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
by: Liu, Jiazheng, et al.
Published: (2025)
by: Liu, Jiazheng, et al.
Published: (2025)
UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026)
by: Li, Boyu, et al.
Published: (2026)
AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback
by: Zhang, Wanpeng, et al.
Published: (2023)
by: Zhang, Wanpeng, et al.
Published: (2023)
DemoGrasp: Universal Dexterous Grasping from a Single Demonstration
by: Yuan, Haoqi, et al.
Published: (2025)
by: Yuan, Haoqi, et al.
Published: (2025)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
OpenT2M: No-frill Motion Generation with Open-source,Large-scale, High-quality Data
by: Cao, Bin, et al.
Published: (2026)
by: Cao, Bin, et al.
Published: (2026)
Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations
by: Zhou, Bohan, et al.
Published: (2024)
by: Zhou, Bohan, et al.
Published: (2024)
Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learning
by: Mao, Chuan, et al.
Published: (2025)
by: Mao, Chuan, et al.
Published: (2025)
DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation
by: Fu, Yuhui, et al.
Published: (2025)
by: Fu, Yuhui, et al.
Published: (2025)
Robust Motion Generation using Part-level Reliable Data from Videos
by: Li, Boyuan, et al.
Published: (2025)
by: Li, Boyuan, et al.
Published: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping
by: Huang, Ziye, et al.
Published: (2024)
by: Huang, Ziye, et al.
Published: (2024)
Cross-Embodiment Dexterous Grasping with Reinforcement Learning
by: Yuan, Haoqi, et al.
Published: (2024)
by: Yuan, Haoqi, et al.
Published: (2024)
QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds
by: Mei, Yuting, et al.
Published: (2024)
by: Mei, Yuting, et al.
Published: (2024)
Tackling Non-Stationarity in Reinforcement Learning via Causal-Origin Representation
by: Zhang, Wanpeng, et al.
Published: (2023)
by: Zhang, Wanpeng, et al.
Published: (2023)
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
SPAFormer: Sequential 3D Part Assembly with Transformers
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models
by: Zhou, Tian-Yi, et al.
Published: (2026)
by: Zhou, Tian-Yi, et al.
Published: (2026)
Creative Agents: Empowering Agents with Imagination for Creative Tasks
by: Cai, Penglin, et al.
Published: (2023)
by: Cai, Penglin, et al.
Published: (2023)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
Latent Action Pretraining from Videos
by: Ye, Seonghyeon, et al.
Published: (2024)
by: Ye, Seonghyeon, et al.
Published: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
by: Xiong, Weimin, et al.
Published: (2026)
by: Xiong, Weimin, et al.
Published: (2026)
Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
by: Zhu, Chuning, et al.
Published: (2025)
by: Zhu, Chuning, et al.
Published: (2025)
Similar Items
-
Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
by: Feng, Yicheng, et al.
Published: (2025) -
Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization
by: Luo, Hao, et al.
Published: (2026) -
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026) -
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026) -
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
by: Wang, Ye, et al.
Published: (2026)