HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Liyang, Ma, Tianxiang, Liu, Jiawei, Li, Bingchuan, Chen, Zhuowei, Liu, Lijie, He, Xu, Li, Gen, He, Qian, Wu, Zhiyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916944740876288
author Chen, Liyang
Ma, Tianxiang
Liu, Jiawei
Li, Bingchuan
Chen, Zhuowei
Liu, Lijie
He, Xu
Li, Gen
He, Qian
Wu, Zhiyong
author_facet Chen, Liyang
Ma, Tianxiang
Liu, Jiawei
Li, Bingchuan
Chen, Zhuowei
Liu, Lijie
He, Xu
Li, Gen
He, Qian
Wu, Zhiyong
contents Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, image, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of training data with paired triplet conditions and the difficulty of collaborating the sub-tasks of subject preservation and audio-visual sync with multimodal inputs. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct a high-quality dataset with diverse and paired text, reference images, and audio. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies. For the subject preservation task, to maintain the prompt following and visual generation abilities of the foundation model, we adopt the minimal-invasive image injection strategy. For the audio-visual sync task, besides the commonly adopted audio cross-attention layer, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multimodal inputs, building on previously acquired capabilities, we progressively incorporate the audio-visual sync task. During inference, for flexible and fine-grained multimodal control, we design a time-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG. Project Page: https://phantom-video.github.io/HuMo.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08519
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Chen, Liyang
Ma, Tianxiang
Liu, Jiawei
Li, Bingchuan
Chen, Zhuowei
Liu, Lijie
He, Xu
Li, Gen
He, Qian
Wu, Zhiyong
Computer Vision and Pattern Recognition
Multimedia
Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, image, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of training data with paired triplet conditions and the difficulty of collaborating the sub-tasks of subject preservation and audio-visual sync with multimodal inputs. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct a high-quality dataset with diverse and paired text, reference images, and audio. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies. For the subject preservation task, to maintain the prompt following and visual generation abilities of the foundation model, we adopt the minimal-invasive image injection strategy. For the audio-visual sync task, besides the commonly adopted audio cross-attention layer, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multimodal inputs, building on previously acquired capabilities, we progressively incorporate the audio-visual sync task. During inference, for flexible and fine-grained multimodal control, we design a time-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG. Project Page: https://phantom-video.github.io/HuMo.
title HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2509.08519