InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhenzhi, Yang, Jiaqi, Jiang, Jianwen, Liang, Chao, Lin, Gaojie, Zheng, Zerong, Yang, Ceyuan, Zhang, Yuan, Gao, Mingyuan, Lin, Dahua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911485250240512
author Wang, Zhenzhi
Yang, Jiaqi
Jiang, Jianwen
Liang, Chao
Lin, Gaojie
Zheng, Zerong
Yang, Ceyuan
Zhang, Yuan
Gao, Mingyuan
Lin, Dahua
author_facet Wang, Zhenzhi
Yang, Jiaqi
Jiang, Jianwen
Liang, Chao
Lin, Gaojie
Zheng, Zerong
Yang, Ceyuan
Zhang, Yuan
Gao, Mingyuan
Lin, Dahua
contents End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios where multiple concepts could appear in the same video with rich human-human interactions and human-object interactions. Such a global assumption prevents precise and per-identity control of multiple concepts including humans and objects, therefore hinders applications. In this work, we discard the single-entity assumption and introduce a novel framework that enforces strong, region-specific binding of conditions from modalities to each identity's spatiotemporal footprint. Given reference images of multiple concepts, our method could automatically infer layout information by leveraging a mask predictor to match appearance cues between the denoised video and each reference appearance. Furthermore, we inject local audio condition into its corresponding region to ensure layout-aligned modality matching in an iterative manner. This design enables the high-quality generation of human dialogue videos between two to three people or video customization from multiple reference images. Empirical results and ablation studies validate the effectiveness of our explicit layout control for multi-modal conditions compared to implicit counterparts and other existing methods. Video demos are available at https://zhenzhiwang.github.io/interacthuman/
format Preprint
id arxiv_https___arxiv_org_abs_2506_09984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
Wang, Zhenzhi
Yang, Jiaqi
Jiang, Jianwen
Liang, Chao
Lin, Gaojie
Zheng, Zerong
Yang, Ceyuan
Zhang, Yuan
Gao, Mingyuan
Lin, Dahua
Computer Vision and Pattern Recognition
Artificial Intelligence
Sound
End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios where multiple concepts could appear in the same video with rich human-human interactions and human-object interactions. Such a global assumption prevents precise and per-identity control of multiple concepts including humans and objects, therefore hinders applications. In this work, we discard the single-entity assumption and introduce a novel framework that enforces strong, region-specific binding of conditions from modalities to each identity's spatiotemporal footprint. Given reference images of multiple concepts, our method could automatically infer layout information by leveraging a mask predictor to match appearance cues between the denoised video and each reference appearance. Furthermore, we inject local audio condition into its corresponding region to ensure layout-aligned modality matching in an iterative manner. This design enables the high-quality generation of human dialogue videos between two to three people or video customization from multiple reference images. Empirical results and ablation studies validate the effectiveness of our explicit layout control for multi-modal conditions compared to implicit counterparts and other existing methods. Video demos are available at https://zhenzhiwang.github.io/interacthuman/
title InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Sound
url https://arxiv.org/abs/2506.09984