VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909900193398784 |
|---|---|
| author | Yin, Shaofeng Ze, Yanjie Yu, Hong-Xing Liu, C. Karen Wu, Jiajun |
| author_facet | Yin, Shaofeng Ze, Yanjie Yu, Hong-Xing Liu, C. Karen Wu, Jiajun |
| contents | Humanoid loco-manipulation in unstructured environments demands tight integration of egocentric perception and whole-body control. However, existing approaches either depend on external motion capture systems or fail to generalize across diverse tasks. We introduce VisualMimic, a visual sim-to-real framework that unifies egocentric vision with hierarchical whole-body control for humanoid robots. VisualMimic combines a task-agnostic low-level keypoint tracker -- trained from human motion data via a teacher-student scheme -- with a task-specific high-level policy that generates keypoint commands from visual and proprioceptive input. To ensure stable training, we inject noise into the low-level policy and clip high-level actions using human motion statistics. VisualMimic enables zero-shot transfer of visuomotor policies trained in simulation to real humanoid robots, accomplishing a wide range of loco-manipulation tasks such as box lifting, pushing, football dribbling, and kicking. Beyond controlled laboratory settings, our policies also generalize robustly to outdoor environments. Videos are available at: https://visualmimic.github.io . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_20322 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation Yin, Shaofeng Ze, Yanjie Yu, Hong-Xing Liu, C. Karen Wu, Jiajun Robotics Computer Vision and Pattern Recognition Machine Learning Humanoid loco-manipulation in unstructured environments demands tight integration of egocentric perception and whole-body control. However, existing approaches either depend on external motion capture systems or fail to generalize across diverse tasks. We introduce VisualMimic, a visual sim-to-real framework that unifies egocentric vision with hierarchical whole-body control for humanoid robots. VisualMimic combines a task-agnostic low-level keypoint tracker -- trained from human motion data via a teacher-student scheme -- with a task-specific high-level policy that generates keypoint commands from visual and proprioceptive input. To ensure stable training, we inject noise into the low-level policy and clip high-level actions using human motion statistics. VisualMimic enables zero-shot transfer of visuomotor policies trained in simulation to real humanoid robots, accomplishing a wide range of loco-manipulation tasks such as box lifting, pushing, football dribbling, and kicking. Beyond controlled laboratory settings, our policies also generalize robustly to outdoor environments. Videos are available at: https://visualmimic.github.io . |
| title | VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation |
| topic | Robotics Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2509.20322 |