Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiahe, Wang, ZiRui, Jia, Feiyu, Chen, Xiao, Niu, Xiaojie, Zeng, Weishuai, Xue, Tianfan, Zhou, Xiaowei, Pang, Jiangmiao, Wang, Jingbo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911706797572096
author Chen, Jiahe
Wang, ZiRui
Jia, Feiyu
Chen, Xiao
Niu, Xiaojie
Zeng, Weishuai
Xue, Tianfan
Zhou, Xiaowei
Pang, Jiangmiao
Wang, Jingbo
author_facet Chen, Jiahe
Wang, ZiRui
Jia, Feiyu
Chen, Xiao
Niu, Xiaojie
Zeng, Weishuai
Xue, Tianfan
Zhou, Xiaowei
Pang, Jiangmiao
Wang, Jingbo
contents Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to their reliance on geometric priors (e.g., explicit CAD models), and \textit{Retargeting Complexity} arising from intensive morphing and morphological mismatch. We propose Imagine2Real, a zero-shot HOI framework for flexible, geometry-free interaction. To resolve misalignment, we formulate robot and object motions as unified 4D point trajectories. To overcome retargeting complexity, our Keypoints Tracker tracks only sparse critical points (base, hands, and object), entirely bypassing the error-amplifying retargeting process. To maintain natural gaits despite these sparse signals, we utilize the latent space of a Behavior Foundation Model (BFM) as the tracker's search domain. Using a progressive training strategy, Imagine2Real learns robust behaviors with simple tracking rewards, enabling zero-shot physical deployment within a motion capture(mocap) system.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22272
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
Chen, Jiahe
Wang, ZiRui
Jia, Feiyu
Chen, Xiao
Niu, Xiaojie
Zeng, Weishuai
Xue, Tianfan
Zhou, Xiaowei
Pang, Jiangmiao
Wang, Jingbo
Robotics
Computer Vision and Pattern Recognition
Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to their reliance on geometric priors (e.g., explicit CAD models), and \textit{Retargeting Complexity} arising from intensive morphing and morphological mismatch. We propose Imagine2Real, a zero-shot HOI framework for flexible, geometry-free interaction. To resolve misalignment, we formulate robot and object motions as unified 4D point trajectories. To overcome retargeting complexity, our Keypoints Tracker tracks only sparse critical points (base, hands, and object), entirely bypassing the error-amplifying retargeting process. To maintain natural gaits despite these sparse signals, we utilize the latent space of a Behavior Foundation Model (BFM) as the tracker's search domain. Using a progressive training strategy, Imagine2Real learns robust behaviors with simple tracking rewards, enabling zero-shot physical deployment within a motion capture(mocap) system.
title Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.22272