Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qixiu, Deng, Yu, Liang, Yaobo, Luo, Lin, Zhou, Lei, Yao, Chengtang, Zeng, Lingqi, Feng, Zhiyuan, Liang, Huizhi, Xu, Sicheng, Zhang, Yizhong, Chen, Xi, Chen, Hao, Sun, Lily, Chen, Dong, Yang, Jiaolong, Guo, Baining |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024)
by: Li, Qixiu, et al.
Published: (2024)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026)
by: Liang, Huizhi, et al.
Published: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026)
by: Xu, Sicheng, et al.
Published: (2026)
UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
by: Wang, Wenbo, et al.
Published: (2024)
by: Wang, Wenbo, et al.
Published: (2024)
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
by: Wang, Wenbo, et al.
Published: (2026)
by: Wang, Wenbo, et al.
Published: (2026)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025)
by: Xu, Sicheng, et al.
Published: (2025)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
by: Feng, ZhiYuan, et al.
Published: (2026)
by: Feng, ZhiYuan, et al.
Published: (2026)
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
by: Xu, Sicheng, et al.
Published: (2024)
by: Xu, Sicheng, et al.
Published: (2024)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
by: Du, Zhiying, et al.
Published: (2025)
by: Du, Zhiying, et al.
Published: (2025)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
by: Song, Zijian, et al.
Published: (2026)
by: Song, Zijian, et al.
Published: (2026)
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
by: Chen, Yixiang, et al.
Published: (2025)
by: Chen, Yixiang, et al.
Published: (2025)
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation
by: Yan, Hongyu, et al.
Published: (2026)
by: Yan, Hongyu, et al.
Published: (2026)
Hyperbolic Multiview Pretraining for Robotic Manipulation
by: Yang, Jin, et al.
Published: (2026)
by: Yang, Jin, et al.
Published: (2026)
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
by: Guo, Shanshan, et al.
Published: (2025)
by: Guo, Shanshan, et al.
Published: (2025)
Structured 3D Latents for Scalable and Versatile 3D Generation
by: Xiang, Jianfeng, et al.
Published: (2024)
by: Xiang, Jianfeng, et al.
Published: (2024)
Revisiting Prompt Pretraining of Vision-Language Models
by: Chen, Zhenyuan, et al.
Published: (2024)
by: Chen, Zhenyuan, et al.
Published: (2024)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
by: Zhang, Yihao, et al.
Published: (2025)
by: Zhang, Yihao, et al.
Published: (2025)
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
by: Song, Zijian, et al.
Published: (2026)
by: Song, Zijian, et al.
Published: (2026)
VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
by: Bian, Jinyue, et al.
Published: (2025)
by: Bian, Jinyue, et al.
Published: (2025)
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025)
by: Zhou, Zhiyuan, et al.
Published: (2025)
LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
by: Wang, Xiaofan, et al.
Published: (2025)
by: Wang, Xiaofan, et al.
Published: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
Human2Robot: Learning Robot Actions from Paired Human-Robot Videos
by: Xie, Sicheng, et al.
Published: (2025)
by: Xie, Sicheng, et al.
Published: (2025)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
by: Xie, Pengzhen, et al.
Published: (2025)
by: Xie, Pengzhen, et al.
Published: (2025)
nicolay-r at SemEval-2024 Task 3: Using Flan-T5 for Reasoning Emotion Cause in Conversations with Chain-of-Thought on Emotion States
by: Rusnachenko, Nicolay, et al.
Published: (2024)
by: Rusnachenko, Nicolay, et al.
Published: (2024)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
by: Zhou, Pengfei, et al.
Published: (2026)
by: Zhou, Pengfei, et al.
Published: (2026)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
by: Yang, Rushuai, et al.
Published: (2025)
by: Yang, Rushuai, et al.
Published: (2025)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
by: Shao, Rui, et al.
Published: (2025)
by: Shao, Rui, et al.
Published: (2025)
Is Diversity All You Need for Scalable Robotic Manipulation?
by: Shi, Modi, et al.
Published: (2025)
by: Shi, Modi, et al.
Published: (2025)
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
by: Pei, Xiaohuan, et al.
Published: (2025)
by: Pei, Xiaohuan, et al.
Published: (2025)
Similar Items
-
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024) -
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026) -
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025) -
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025) -
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026)