iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Zhelun, Wu, Chenming, Zhou, Junsheng, Zhao, Chen, Wang, Kaisiyuan, Zhou, Hang, Li, Yingying, Feng, Haocheng, He, Wei, Wang, Jingdong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912431812378624
author Shen, Zhelun
Wu, Chenming
Zhou, Junsheng
Zhao, Chen
Wang, Kaisiyuan
Zhou, Hang
Li, Yingying
Feng, Haocheng
He, Wei
Wang, Jingdong
author_facet Shen, Zhelun
Wu, Chenming
Zhou, Junsheng
Zhao, Chen
Wang, Kaisiyuan
Zhou, Hang
Li, Yingying
Feng, Haocheng
He, Wei
Wang, Jingdong
contents Digital human video generation is gaining traction in fields like education and e-commerce, driven by advancements in head-body animation and lip-syncing technologies. However, realistic Hand-Object Interaction (HOI) - the complex dynamics between human hands and objects - continues to pose challenges. Generating natural and believable HOI reenactments is difficult due to issues such as occlusion between hands and objects, variations in object shapes and orientations, and the necessity for precise physical interactions, and importantly, the ability to generalize to unseen humans and objects. This paper presents a novel framework iDiT-HOI that enables in-the-wild HOI reenactment generation. Specifically, we propose a unified inpainting-based token process method, called Inp-TPU, with a two-stage video diffusion transformer (DiT) model. The first stage generates a key frame by inserting the designated object into the hand region, providing a reference for subsequent frames. The second stage ensures temporal coherence and fluidity in hand-object interactions. The key contribution of our method is to reuse the pretrained model's context perception capabilities without introducing additional parameters, enabling strong generalization to unseen objects and scenarios, and our proposed paradigm naturally supports long video generation. Comprehensive evaluations demonstrate that our approach outperforms existing methods, particularly in challenging real-world scenes, offering enhanced realism and more seamless hand-object interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
Shen, Zhelun
Wu, Chenming
Zhou, Junsheng
Zhao, Chen
Wang, Kaisiyuan
Zhou, Hang
Li, Yingying
Feng, Haocheng
He, Wei
Wang, Jingdong
Graphics
Computer Vision and Pattern Recognition
Digital human video generation is gaining traction in fields like education and e-commerce, driven by advancements in head-body animation and lip-syncing technologies. However, realistic Hand-Object Interaction (HOI) - the complex dynamics between human hands and objects - continues to pose challenges. Generating natural and believable HOI reenactments is difficult due to issues such as occlusion between hands and objects, variations in object shapes and orientations, and the necessity for precise physical interactions, and importantly, the ability to generalize to unseen humans and objects. This paper presents a novel framework iDiT-HOI that enables in-the-wild HOI reenactment generation. Specifically, we propose a unified inpainting-based token process method, called Inp-TPU, with a two-stage video diffusion transformer (DiT) model. The first stage generates a key frame by inserting the designated object into the hand region, providing a reference for subsequent frames. The second stage ensures temporal coherence and fluidity in hand-object interactions. The key contribution of our method is to reuse the pretrained model's context perception capabilities without introducing additional parameters, enabling strong generalization to unseen objects and scenarios, and our proposed paradigm naturally supports long video generation. Comprehensive evaluations demonstrate that our approach outperforms existing methods, particularly in challenging real-world scenes, offering enhanced realism and more seamless hand-object interactions.
title iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
topic Graphics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.12847