Ego-Vision World Model for Humanoid Contact Planning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917323869257728 |
|---|---|
| author | Liu, Hang Gao, Yuman Teng, Sangli Chi, Yufeng Shao, Yakun Sophia Li, Zhongyu Ghaffari, Maani Sreenath, Koushil |
| author_facet | Liu, Hang Gao, Yuman Teng, Sangli Chi, Yufeng Shao, Yakun Sophia Li, Zhongyu Ghaffari, Maani Sreenath, Koushil |
| contents | Enabling humanoid robots to exploit physical contact, rather than simply avoid collisions, is crucial for autonomy in unstructured environments. Traditional optimization-based planners struggle with contact complexity, while on-policy reinforcement learning (RL) is sample-inefficient and has limited multi-task ability. We propose a framework combining a learned world model with sampling-based Model Predictive Control (MPC), trained on a demonstration-free offline dataset to predict future outcomes in a compressed latent space. To address sparse contact rewards and sensor noise, the MPC uses a learned surrogate value function for dense, robust planning. Our single, scalable model supports contact-aware tasks, including wall support after perturbation, blocking incoming objects, and traversing height-limited arches, with improved sample efficiency and multi-task capability over on-policy RL. Deployed on a physical humanoid, our system achieves robust, real-time contact planning from proprioception and ego-centric depth images. Code and dataset are available at our website: https://ego-vcp.github.io/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_11682 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Ego-Vision World Model for Humanoid Contact Planning Liu, Hang Gao, Yuman Teng, Sangli Chi, Yufeng Shao, Yakun Sophia Li, Zhongyu Ghaffari, Maani Sreenath, Koushil Robotics Artificial Intelligence Systems and Control Enabling humanoid robots to exploit physical contact, rather than simply avoid collisions, is crucial for autonomy in unstructured environments. Traditional optimization-based planners struggle with contact complexity, while on-policy reinforcement learning (RL) is sample-inefficient and has limited multi-task ability. We propose a framework combining a learned world model with sampling-based Model Predictive Control (MPC), trained on a demonstration-free offline dataset to predict future outcomes in a compressed latent space. To address sparse contact rewards and sensor noise, the MPC uses a learned surrogate value function for dense, robust planning. Our single, scalable model supports contact-aware tasks, including wall support after perturbation, blocking incoming objects, and traversing height-limited arches, with improved sample efficiency and multi-task capability over on-policy RL. Deployed on a physical humanoid, our system achieves robust, real-time contact planning from proprioception and ego-centric depth images. Code and dataset are available at our website: https://ego-vcp.github.io/ |
| title | Ego-Vision World Model for Humanoid Contact Planning |
| topic | Robotics Artificial Intelligence Systems and Control |
| url | https://arxiv.org/abs/2510.11682 |