iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Zhelun, Wu, Chenming, Zhou, Junsheng, Zhao, Chen, Wang, Kaisiyuan, Zhou, Hang, Li, Yingying, Feng, Haocheng, He, Wei, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
by: Fan, Yingying, et al.
Published: (2025)
by: Fan, Yingying, et al.
Published: (2025)
TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
by: Huang, Xuan, et al.
Published: (2026)
by: Huang, Xuan, et al.
Published: (2026)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
by: Tong, Jinguang, et al.
Published: (2026)
by: Tong, Jinguang, et al.
Published: (2026)
Splatter-360: Generalizable 360$^{\circ}$ Gaussian Splatting for Wide-baseline Panoramic Images
by: Chen, Zheng, et al.
Published: (2024)
by: Chen, Zheng, et al.
Published: (2024)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
NeuS-PIR: Learning Relightable Neural Surface using Pre-Integrated Rendering
by: Mao, Shi, et al.
Published: (2023)
by: Mao, Shi, et al.
Published: (2023)
AnyAct: Towards Human Reenactment of Character Motion From Video
by: Chen, Liuhan, et al.
Published: (2026)
by: Chen, Liuhan, et al.
Published: (2026)
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
LuxDiT: Lighting Estimation with Video Diffusion Transformer
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
Surfel-based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction from Monocular Video
by: Zhao, Yiqun, et al.
Published: (2024)
by: Zhao, Yiqun, et al.
Published: (2024)
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
by: Li, Lei, et al.
Published: (2025)
by: Li, Lei, et al.
Published: (2025)
DiffPop: Plausibility‐Guided Object Placement Diffusion for Image Composition
by: Jiacheng Liu, et al.
Published: (2024)
by: Jiacheng Liu, et al.
Published: (2024)
Free Your Hands: Lightweight Turntable-Based Object Capture Pipeline
by: Fan, Jiahui, et al.
Published: (2025)
by: Fan, Jiahui, et al.
Published: (2025)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
From Rigging to Waving: 3D-Guided Diffusion for Natural Animation of Hand-Drawn Characters
by: Zhou, Jie, et al.
Published: (2025)
by: Zhou, Jie, et al.
Published: (2025)
Learning Generalizable Hand-Object Tracking from Synthetic Demonstrations
by: Wang, Yinhuai, et al.
Published: (2025)
by: Wang, Yinhuai, et al.
Published: (2025)
BOOTPLACE: Bootstrapped Object Placement with Detection Transformers
by: Zhou, Hang, et al.
Published: (2025)
by: Zhou, Hang, et al.
Published: (2025)
VideoMat: Extracting PBR Materials from Video Diffusion Models
by: J. Munkberg, et al.
Published: (2025)
by: J. Munkberg, et al.
Published: (2025)
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
by: Peng, Weikun, et al.
Published: (2025)
by: Peng, Weikun, et al.
Published: (2025)
Infusion: Internal Diffusion for Inpainting of Dynamic Textures and Complex Motion
by: N. Cherel, et al.
Published: (2025)
by: N. Cherel, et al.
Published: (2025)
TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size
by: Lionar, Stefan, et al.
Published: (2026)
by: Lionar, Stefan, et al.
Published: (2026)
RI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion Priors
by: Paliwal, Avinash, et al.
Published: (2025)
by: Paliwal, Avinash, et al.
Published: (2025)
Hand-Object Interaction Controller (HOIC): Deep Reinforcement Learning for Reconstructing Interactions with Physics
by: Hu, Haoyu, et al.
Published: (2024)
by: Hu, Haoyu, et al.
Published: (2024)
GeneOH Diffusion: Towards Generalizable Hand-Object Interaction Denoising via Denoising Diffusion
by: Liu, Xueyi, et al.
Published: (2024)
by: Liu, Xueyi, et al.
Published: (2024)
TransVDM: Motion-Constrained Video Diffusion Model for Transparent Video Synthesis
by: Li, Menghao, et al.
Published: (2025)
by: Li, Menghao, et al.
Published: (2025)
NePHIM: A Neural Physics-Based Head-Hand Interaction Model
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
In-Context LoRA for Diffusion Transformers
by: Huang, Lianghua, et al.
Published: (2024)
by: Huang, Lianghua, et al.
Published: (2024)
HOI-Brain: a novel multi-channel transformers framework for brain disorder diagnosis by accurately extracting signed higher-order interactions from fMRI
by: Zhao, Dengyi, et al.
Published: (2025)
by: Zhao, Dengyi, et al.
Published: (2025)
Synchronize Dual Hands for Physics-Based Dexterous Guitar Playing
by: Xu, Pei, et al.
Published: (2024)
by: Xu, Pei, et al.
Published: (2024)
VOODOO XP: Expressive One-Shot Head Reenactment for VR Telepresence
by: Tran, Phong, et al.
Published: (2024)
by: Tran, Phong, et al.
Published: (2024)
Reenact Anything: Semantic Video Motion Transfer Using Motion-Textual Inversion
by: Kansy, Manuel, et al.
Published: (2024)
by: Kansy, Manuel, et al.
Published: (2024)
FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise
by: Yuan, Yunlong, et al.
Published: (2025)
by: Yuan, Yunlong, et al.
Published: (2025)
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
by: Guan, Jiazhi, et al.
Published: (2026)
by: Guan, Jiazhi, et al.
Published: (2026)
VideoMat: Extracting PBR Materials from Video Diffusion Models
by: Munkberg, Jacob, et al.
Published: (2025)
by: Munkberg, Jacob, et al.
Published: (2025)
Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects
by: Barda, Amir, et al.
Published: (2024)
by: Barda, Amir, et al.
Published: (2024)
S ee 4D: Pose‐Free 4D Generation via Auto‐Regressive Video Inpainting
by: Dongyue Lu, et al.
Published: (2026)
by: Dongyue Lu, et al.
Published: (2026)
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
by: Christen, Sammy, et al.
Published: (2024)
by: Christen, Sammy, et al.
Published: (2024)
Similar Items
-
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
by: Fan, Yingying, et al.
Published: (2025) -
TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
by: Guan, Jiazhi, et al.
Published: (2024) -
GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
by: Huang, Xuan, et al.
Published: (2026) -
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025) -
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
by: Tong, Jinguang, et al.
Published: (2026)