DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pang, Yatian, Zhu, Bin, Lin, Bin, Zheng, Mingzhe, Tay, Francis E. H., Lim, Ser-Nam, Yang, Harry, Yuan, Li
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916501154430976
author Pang, Yatian
Zhu, Bin
Lin, Bin
Zheng, Mingzhe
Tay, Francis E. H.
Lim, Ser-Nam
Yang, Harry
Yuan, Li
author_facet Pang, Yatian
Zhu, Bin
Lin, Bin
Zheng, Mingzhe
Tay, Francis E. H.
Lim, Ser-Nam
Yang, Harry
Yuan, Li
contents In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and user-friendly manner. Concretely, baseline methods relying on only 2D pose guidance lack the cues of 3D information, leading to suboptimal results, while methods using 3D representation as guidance achieve higher quality but involve a cumbersome and time-intensive process. To address these limitations, DreamDance enriches 3D geometry cues from 2D poses by introducing an efficient diffusion model, enabling high-quality human image animation with various guidance. Our key insight is that human images naturally exhibit multiple levels of correlation, progressing from coarse skeleton poses to fine-grained geometry cues, and further from these geometry cues to explicit appearance details. Capturing such correlations could enrich the guidance signals, facilitating intra-frame coherency and inter-frame consistency. Specifically, we construct the TikTok-Dance5K dataset, comprising 5K high-quality dance videos with detailed frame annotations, including human pose, depth, and normal maps. Next, we introduce a Mutually Aligned Geometry Diffusion Model to generate fine-grained depth and normal maps for enriched guidance. Finally, a Cross-domain Controller incorporates multi-level guidance to animate human images effectively with a video diffusion model. Extensive experiments demonstrate that our method achieves state-of-the-art performance in animating human images.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00397
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
Pang, Yatian
Zhu, Bin
Lin, Bin
Zheng, Mingzhe
Tay, Francis E. H.
Lim, Ser-Nam
Yang, Harry
Yuan, Li
Computer Vision and Pattern Recognition
In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and user-friendly manner. Concretely, baseline methods relying on only 2D pose guidance lack the cues of 3D information, leading to suboptimal results, while methods using 3D representation as guidance achieve higher quality but involve a cumbersome and time-intensive process. To address these limitations, DreamDance enriches 3D geometry cues from 2D poses by introducing an efficient diffusion model, enabling high-quality human image animation with various guidance. Our key insight is that human images naturally exhibit multiple levels of correlation, progressing from coarse skeleton poses to fine-grained geometry cues, and further from these geometry cues to explicit appearance details. Capturing such correlations could enrich the guidance signals, facilitating intra-frame coherency and inter-frame consistency. Specifically, we construct the TikTok-Dance5K dataset, comprising 5K high-quality dance videos with detailed frame annotations, including human pose, depth, and normal maps. Next, we introduce a Mutually Aligned Geometry Diffusion Model to generate fine-grained depth and normal maps for enriched guidance. Finally, a Cross-domain Controller incorporates multi-level guidance to animate human images effectively with a video diffusion model. Extensive experiments demonstrate that our method achieves state-of-the-art performance in animating human images.
title DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00397