Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ronghui, Zhang, YuXiang, Zhang, Yachao, Zhang, Hongwen, Guo, Jie, Zhang, Yan, Liu, Yebin, Li, Xiu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909176053104640
author Li, Ronghui
Zhang, YuXiang
Zhang, Yachao
Zhang, Hongwen
Guo, Jie
Zhang, Yan
Liu, Yebin
Li, Xiu
author_facet Li, Ronghui
Zhang, YuXiang
Zhang, Yachao
Zhang, Hongwen
Guo, Jie
Zhang, Yan
Liu, Yebin
Li, Xiu
contents We propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations between two diffusion models. The first stage is global diffusion, which focuses on comprehending the coarse-level music-dance correlation and production characteristic dance primitives. In contrast, the second-stage is the local diffusion, which parallelly generates detailed motion sequences under the guidance of the dance primitives and choreographic rules. In addition, we propose a Foot Refine Block to optimize the contact between the feet and the ground, enhancing the physical realism of the motion. Our approach can parallelly generate dance sequences of extremely long length, striking a balance between global choreographic patterns and local motion quality and expressiveness. Extensive experiments validate the efficacy of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10518
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
Li, Ronghui
Zhang, YuXiang
Zhang, Yachao
Zhang, Hongwen
Guo, Jie
Zhang, Yan
Liu, Yebin
Li, Xiu
Computer Vision and Pattern Recognition
Graphics
Sound
Audio and Speech Processing
We propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations between two diffusion models. The first stage is global diffusion, which focuses on comprehending the coarse-level music-dance correlation and production characteristic dance primitives. In contrast, the second-stage is the local diffusion, which parallelly generates detailed motion sequences under the guidance of the dance primitives and choreographic rules. In addition, we propose a Foot Refine Block to optimize the contact between the feet and the ground, enhancing the physical realism of the motion. Our approach can parallelly generate dance sequences of extremely long length, striking a balance between global choreographic patterns and local motion quality and expressiveness. Extensive experiments validate the efficacy of our method.
title Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
topic Computer Vision and Pattern Recognition
Graphics
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2403.10518