AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guan, Jiazhi, Wang, Kaisiyuan, Xu, Zhiliang, Yang, Quanwei, Sun, Yasheng, He, Shengyi, Liang, Borong, Cao, Yukang, Li, Yingying, Feng, Haocheng, Ding, Errui, Wang, Jingdong, Zhao, Youjian, Zhou, Hang, Liu, Ziwei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909643637260288
author Guan, Jiazhi
Wang, Kaisiyuan
Xu, Zhiliang
Yang, Quanwei
Sun, Yasheng
He, Shengyi
Liang, Borong
Cao, Yukang
Li, Yingying
Feng, Haocheng
Ding, Errui
Wang, Jingdong
Zhao, Youjian
Zhou, Hang
Liu, Ziwei
author_facet Guan, Jiazhi
Wang, Kaisiyuan
Xu, Zhiliang
Yang, Quanwei
Sun, Yasheng
He, Shengyi
Liang, Borong
Cao, Yukang
Li, Yingying
Feng, Haocheng
Ding, Errui
Wang, Jingdong
Zhao, Youjian
Zhou, Hang
Liu, Ziwei
contents Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speech gestures w.r.t. given audio. In this work, we propose AudCast, a generalized audio-driven human video generation framework adopting a cascade Diffusion-Transformers (DiTs) paradigm, which synthesizes holistic human videos based on a reference image and a given audio. 1) Firstly, an audio-conditioned Holistic Human DiT architecture is proposed to directly drive the movements of any human body with vivid gesture dynamics. 2) Then to enhance hand and face details that are well-knownly difficult to handle, a Regional Refinement DiT leverages regional 3D fitting as the bridge to reform the signals, producing the final results. Extensive experiments demonstrate that our framework generates high-fidelity audio-driven holistic human videos with temporal coherence and fine facial and hand details. Resources can be found at https://guanjz20.github.io/projects/AudCast.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19824
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
Guan, Jiazhi
Wang, Kaisiyuan
Xu, Zhiliang
Yang, Quanwei
Sun, Yasheng
He, Shengyi
Liang, Borong
Cao, Yukang
Li, Yingying
Feng, Haocheng
Ding, Errui
Wang, Jingdong
Zhao, Youjian
Zhou, Hang
Liu, Ziwei
Computer Vision and Pattern Recognition
Graphics
Multimedia
Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speech gestures w.r.t. given audio. In this work, we propose AudCast, a generalized audio-driven human video generation framework adopting a cascade Diffusion-Transformers (DiTs) paradigm, which synthesizes holistic human videos based on a reference image and a given audio. 1) Firstly, an audio-conditioned Holistic Human DiT architecture is proposed to directly drive the movements of any human body with vivid gesture dynamics. 2) Then to enhance hand and face details that are well-knownly difficult to handle, a Regional Refinement DiT leverages regional 3D fitting as the bridge to reform the signals, producing the final results. Extensive experiments demonstrate that our framework generates high-fidelity audio-driven holistic human videos with temporal coherence and fine facial and hand details. Resources can be found at https://guanjz20.github.io/projects/AudCast.
title AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
topic Computer Vision and Pattern Recognition
Graphics
Multimedia
url https://arxiv.org/abs/2503.19824