Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911706797572096 |
|---|---|
| author | Chen, Jiahe Wang, ZiRui Jia, Feiyu Chen, Xiao Niu, Xiaojie Zeng, Weishuai Xue, Tianfan Zhou, Xiaowei Pang, Jiangmiao Wang, Jingbo |
| author_facet | Chen, Jiahe Wang, ZiRui Jia, Feiyu Chen, Xiao Niu, Xiaojie Zeng, Weishuai Xue, Tianfan Zhou, Xiaowei Pang, Jiangmiao Wang, Jingbo |
| contents | Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to their reliance on geometric priors (e.g., explicit CAD models), and \textit{Retargeting Complexity} arising from intensive morphing and morphological mismatch. We propose Imagine2Real, a zero-shot HOI framework for flexible, geometry-free interaction. To resolve misalignment, we formulate robot and object motions as unified 4D point trajectories. To overcome retargeting complexity, our Keypoints Tracker tracks only sparse critical points (base, hands, and object), entirely bypassing the error-amplifying retargeting process. To maintain natural gaits despite these sparse signals, we utilize the latent space of a Behavior Foundation Model (BFM) as the tracker's search domain. Using a progressive training strategy, Imagine2Real learns robust behaviors with simple tracking rewards, enabling zero-shot physical deployment within a motion capture(mocap) system. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_22272 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors Chen, Jiahe Wang, ZiRui Jia, Feiyu Chen, Xiao Niu, Xiaojie Zeng, Weishuai Xue, Tianfan Zhou, Xiaowei Pang, Jiangmiao Wang, Jingbo Robotics Computer Vision and Pattern Recognition Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to their reliance on geometric priors (e.g., explicit CAD models), and \textit{Retargeting Complexity} arising from intensive morphing and morphological mismatch. We propose Imagine2Real, a zero-shot HOI framework for flexible, geometry-free interaction. To resolve misalignment, we formulate robot and object motions as unified 4D point trajectories. To overcome retargeting complexity, our Keypoints Tracker tracks only sparse critical points (base, hands, and object), entirely bypassing the error-amplifying retargeting process. To maintain natural gaits despite these sparse signals, we utilize the latent space of a Behavior Foundation Model (BFM) as the tracker's search domain. Using a progressive training strategy, Imagine2Real learns robust behaviors with simple tracking rewards, enabling zero-shot physical deployment within a motion capture(mocap) system. |
| title | Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.22272 |