DGFM: Full Body Dance Generation Driven by Music Foundation Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Xinran, Feng, Zhenhua, Kanojia, Diptesh, Wang, Wenwu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912250512539648
author Liu, Xinran
Feng, Zhenhua
Kanojia, Diptesh
Wang, Wenwu
author_facet Liu, Xinran
Feng, Zhenhua
Kanojia, Diptesh
Wang, Wenwu
contents In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based method that generates dance movements conditioned on text and music. Our approach extracts music features by combining high-level features obtained by music foundation model with hand-crafted features, thereby enhancing the quality of generated dance sequences. This method effectively leverages the advantages of high-level semantic information and low-level temporal details to improve the model's capability in music feature understanding. To show the merits of the proposed method, we compare it with four music foundation models and two sets of hand-crafted music features. The results demonstrate that our method obtains the most realistic dance sequences and achieves the best match with the input music.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20176
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DGFM: Full Body Dance Generation Driven by Music Foundation Models
Liu, Xinran
Feng, Zhenhua
Kanojia, Diptesh
Wang, Wenwu
Sound
Graphics
Audio and Speech Processing
In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based method that generates dance movements conditioned on text and music. Our approach extracts music features by combining high-level features obtained by music foundation model with hand-crafted features, thereby enhancing the quality of generated dance sequences. This method effectively leverages the advantages of high-level semantic information and low-level temporal details to improve the model's capability in music feature understanding. To show the merits of the proposed method, we compare it with four music foundation models and two sets of hand-crafted music features. The results demonstrate that our method obtains the most realistic dance sequences and achieves the best match with the input music.
title DGFM: Full Body Dance Generation Driven by Music Foundation Models
topic Sound
Graphics
Audio and Speech Processing
url https://arxiv.org/abs/2502.20176