HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Ziyao, Zhou, Zixiang, Cao, Juan, Ma, Yifeng, Chen, Yi, Rao, Zejing, Xu, Zhiyong, Wang, Hongmei, Lin, Qin, Zhou, Yuan, Lu, Qinglin, Tang, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912422661455872
author Huang, Ziyao
Zhou, Zixiang
Cao, Juan
Ma, Yifeng
Chen, Yi
Rao, Zejing
Xu, Zhiyong
Wang, Hongmei
Lin, Qin
Zhou, Yuan
Lu, Qinglin
Tang, Fan
author_facet Huang, Ziyao
Zhou, Zixiang
Cao, Juan
Ma, Yifeng
Chen, Yi
Rao, Zejing
Xu, Zhiyong
Wang, Hongmei
Lin, Qin
Zhou, Yuan
Lu, Qinglin
Tang, Fan
contents To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce HunyuanVideo-HOMA, a weakly conditioned multimodal-driven framework. HunyuanVideo-HOMA enhances controllability and reduces dependency on precise inputs through sparse, decoupled motion guidance. It encodes appearance and motion signals into the dual input space of a multimodal diffusion transformer (MMDiT), fusing them within a shared context space to synthesize temporally consistent and physically plausible interactions. To optimize training, we integrate a parameter-space HOI adapter initialized from pretrained MMDiT weights, preserving prior knowledge while enabling efficient adaptation, and a facial cross-attention adapter for anatomically accurate audio-driven lip synchronization. Extensive experiments confirm state-of-the-art performance in interaction naturalness and generalization under weak supervision. Finally, HunyuanVideo-HOMA demonstrates versatility in text-conditioned generation and interactive object manipulation, supported by a user-friendly demo interface. The project page is at https://anonymous.4open.science/w/homa-page-0FBE/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
Huang, Ziyao
Zhou, Zixiang
Cao, Juan
Ma, Yifeng
Chen, Yi
Rao, Zejing
Xu, Zhiyong
Wang, Hongmei
Lin, Qin
Zhou, Yuan
Lu, Qinglin
Tang, Fan
Computer Vision and Pattern Recognition
To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce HunyuanVideo-HOMA, a weakly conditioned multimodal-driven framework. HunyuanVideo-HOMA enhances controllability and reduces dependency on precise inputs through sparse, decoupled motion guidance. It encodes appearance and motion signals into the dual input space of a multimodal diffusion transformer (MMDiT), fusing them within a shared context space to synthesize temporally consistent and physically plausible interactions. To optimize training, we integrate a parameter-space HOI adapter initialized from pretrained MMDiT weights, preserving prior knowledge while enabling efficient adaptation, and a facial cross-attention adapter for anatomically accurate audio-driven lip synchronization. Extensive experiments confirm state-of-the-art performance in interaction naturalness and generalization under weak supervision. Finally, HunyuanVideo-HOMA demonstrates versatility in text-conditioned generation and interactive object manipulation, supported by a user-friendly demo interface. The project page is at https://anonymous.4open.science/w/homa-page-0FBE/.
title HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08797