Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yicheng, Zhang, Wanpeng, Wang, Ye, Luo, Hao, Yuan, Haoqi, Zheng, Sipeng, Lu, Zongqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Conservative Offline Robot Policy Learning via Posterior-Transition Reweighting
by: Zhang, Wanpeng, et al.
Published: (2026)
by: Zhang, Wanpeng, et al.
Published: (2026)
Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
by: Xu, Haiweng, et al.
Published: (2026)
by: Xu, Haiweng, et al.
Published: (2026)
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations
by: Zhou, Bohan, et al.
Published: (2024)
by: Zhou, Bohan, et al.
Published: (2024)
UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026)
by: Li, Boyu, et al.
Published: (2026)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
DemoGrasp: Universal Dexterous Grasping from a Single Demonstration
by: Yuan, Haoqi, et al.
Published: (2025)
by: Yuan, Haoqi, et al.
Published: (2025)
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping
by: Huang, Ziye, et al.
Published: (2024)
by: Huang, Ziye, et al.
Published: (2024)
Cross-Embodiment Dexterous Grasping with Reinforcement Learning
by: Yuan, Haoqi, et al.
Published: (2024)
by: Yuan, Haoqi, et al.
Published: (2024)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
PhysiFlow: Physics-Aware Humanoid Whole-Body VLA via Multi-Brain Latent Flow Matching and Robust Tracking
by: Qin, Weikai, et al.
Published: (2026)
by: Qin, Weikai, et al.
Published: (2026)
Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learning
by: Mao, Chuan, et al.
Published: (2025)
by: Mao, Chuan, et al.
Published: (2025)
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
by: Yuan, Haoqi, et al.
Published: (2025)
by: Yuan, Haoqi, et al.
Published: (2025)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)
by: Zheng, Sipeng, et al.
Published: (2024)
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
by: Yue, Junpeng, et al.
Published: (2025)
by: Yue, Junpeng, et al.
Published: (2025)
DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation
by: Fu, Yuhui, et al.
Published: (2025)
by: Fu, Yuhui, et al.
Published: (2025)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
Visual Robotic Manipulation with Depth-Aware Pretraining
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid
by: Xu, Xinyu, et al.
Published: (2024)
by: Xu, Xinyu, et al.
Published: (2024)
Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA
by: Tang, Tutian, et al.
Published: (2026)
by: Tang, Tutian, et al.
Published: (2026)
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
by: Miao, Shangchen, et al.
Published: (2026)
by: Miao, Shangchen, et al.
Published: (2026)
DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
by: Pan, Xu, et al.
Published: (2026)
by: Pan, Xu, et al.
Published: (2026)
DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
by: Li, Boyu, et al.
Published: (2025)
by: Li, Boyu, et al.
Published: (2025)
AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback
by: Zhang, Wanpeng, et al.
Published: (2023)
by: Zhang, Wanpeng, et al.
Published: (2023)
QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds
by: Mei, Yuting, et al.
Published: (2024)
by: Mei, Yuting, et al.
Published: (2024)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
General Humanoid Whole-Body Control via Pretraining and Fast Adaptation
by: Wang, Zepeng, et al.
Published: (2026)
by: Wang, Zepeng, et al.
Published: (2026)
Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
by: Li, Boyu, et al.
Published: (2025)
by: Li, Boyu, et al.
Published: (2025)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
by: Zhang, Yichi, et al.
Published: (2026)
by: Zhang, Yichi, et al.
Published: (2026)
Similar Items
-
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026) -
DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
by: Zhang, Wanpeng, et al.
Published: (2025) -
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025) -
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
by: Wang, Ye, et al.
Published: (2026) -
Conservative Offline Robot Policy Learning via Posterior-Transition Reweighting
by: Zhang, Wanpeng, et al.
Published: (2026)