InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915448504713216 |
|---|---|
| author | Lin, Yukang Hong, Yan Xu, Zunnan Li, Xindi Xu, Chao Song, Chuanbiao Li, Ronghui Chen, Haoxing Lan, Jun Zhu, Huijia Wang, Weiqiang Zhang, Jianfu Li, Xiu |
| author_facet | Lin, Yukang Hong, Yan Xu, Zunnan Li, Xindi Xu, Chao Song, Chuanbiao Li, Ronghui Chen, Haoxing Lan, Jun Zhu, Huijia Wang, Weiqiang Zhang, Jianfu Li, Xiu |
| contents | Recent video generation research has focused heavily on isolated actions, leaving interactive motions-such as hand-face interactions-largely unexamined. These interactions are essential for emerging biometric authentication systems, which rely on interactive motion-based anti-spoofing approaches. From a security perspective, there is a growing need for large-scale, high-quality interactive videos to train and strengthen authentication models. In this work, we introduce a novel paradigm for animating realistic hand-face interactions. Our approach simultaneously learns spatio-temporal contact dynamics and biomechanically plausible deformation effects, enabling natural interactions where hand movements induce anatomically accurate facial deformations while maintaining collision-free contact. To facilitate this research, we present InterHF, a large-scale hand-face interaction dataset featuring 18 interaction patterns and 90,000 annotated videos. Additionally, we propose InterAnimate, a region-aware diffusion model designed specifically for interaction animation. InterAnimate leverages learnable spatial and temporal latents to effectively capture dynamic interaction priors and integrates a region-aware interaction mechanism that injects these priors into the denoising process. To the best of our knowledge, this work represents the first large-scale effort to systematically study human hand-face interactions. Qualitative and quantitative results show InterAnimate produces highly realistic animations, setting a new benchmark. Code and data will be made public to advance research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_10905 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation Lin, Yukang Hong, Yan Xu, Zunnan Li, Xindi Xu, Chao Song, Chuanbiao Li, Ronghui Chen, Haoxing Lan, Jun Zhu, Huijia Wang, Weiqiang Zhang, Jianfu Li, Xiu Computer Vision and Pattern Recognition Human-Computer Interaction Recent video generation research has focused heavily on isolated actions, leaving interactive motions-such as hand-face interactions-largely unexamined. These interactions are essential for emerging biometric authentication systems, which rely on interactive motion-based anti-spoofing approaches. From a security perspective, there is a growing need for large-scale, high-quality interactive videos to train and strengthen authentication models. In this work, we introduce a novel paradigm for animating realistic hand-face interactions. Our approach simultaneously learns spatio-temporal contact dynamics and biomechanically plausible deformation effects, enabling natural interactions where hand movements induce anatomically accurate facial deformations while maintaining collision-free contact. To facilitate this research, we present InterHF, a large-scale hand-face interaction dataset featuring 18 interaction patterns and 90,000 annotated videos. Additionally, we propose InterAnimate, a region-aware diffusion model designed specifically for interaction animation. InterAnimate leverages learnable spatial and temporal latents to effectively capture dynamic interaction priors and integrates a region-aware interaction mechanism that injects these priors into the denoising process. To the best of our knowledge, this work represents the first large-scale effort to systematically study human hand-face interactions. Qualitative and quantitative results show InterAnimate produces highly realistic animations, setting a new benchmark. Code and data will be made public to advance research. |
| title | InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation |
| topic | Computer Vision and Pattern Recognition Human-Computer Interaction |
| url | https://arxiv.org/abs/2504.10905 |