MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Kaixing, Tang, Xulong, Peng, Ziqiao, Hu, Yuxuan, Zhang, Xiangyue, Wang, Puwei, Liu, Hongyan, He, Jun, Fan, Zhaoxin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912993891057664
author Yang, Kaixing
Tang, Xulong
Peng, Ziqiao
Hu, Yuxuan
Zhang, Xiangyue
Wang, Puwei
Liu, Hongyan
He, Jun
Fan, Zhaoxin
author_facet Yang, Kaixing
Tang, Xulong
Peng, Ziqiao
Hu, Yuxuan
Zhang, Xiangyue
Wang, Puwei
Liu, Hongyan
He, Jun
Fan, Zhaoxin
contents Music-to-dance generation represents a challenging yet pivotal task at the intersection of choreography, virtual reality, and creative content generation. Despite its significance, existing methods face substantial limitation in achieving choreographic consistency. To address the challenge, we propose MatchDance, a novel framework for music-to-dance generation that constructs a latent representation to enhance choreographic consistency. MatchDance employs a two-stage design: (1) a Kinematic-Dynamic-based Quantization Stage (KDQS), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) with kinematic-dynamic constraints and reconstructs them with high fidelity, and (2) a Hybrid Music-to-Dance Generation Stage(HMDGS), which uses a Mamba-Transformer hybrid architecture to map music into the latent representation, followed by the KDQS decoder to generate 3D dance motions. Additionally, a music-dance retrieval framework and comprehensive metrics are introduced for evaluation. Extensive experiments on the FineDance dataset demonstrate state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
Yang, Kaixing
Tang, Xulong
Peng, Ziqiao
Hu, Yuxuan
Zhang, Xiangyue
Wang, Puwei
Liu, Hongyan
He, Jun
Fan, Zhaoxin
Sound
Graphics
Multimedia
Audio and Speech Processing
Music-to-dance generation represents a challenging yet pivotal task at the intersection of choreography, virtual reality, and creative content generation. Despite its significance, existing methods face substantial limitation in achieving choreographic consistency. To address the challenge, we propose MatchDance, a novel framework for music-to-dance generation that constructs a latent representation to enhance choreographic consistency. MatchDance employs a two-stage design: (1) a Kinematic-Dynamic-based Quantization Stage (KDQS), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) with kinematic-dynamic constraints and reconstructs them with high fidelity, and (2) a Hybrid Music-to-Dance Generation Stage(HMDGS), which uses a Mamba-Transformer hybrid architecture to map music into the latent representation, followed by the KDQS decoder to generate 3D dance motions. Additionally, a music-dance retrieval framework and comprehensive metrics are introduced for evaluation. Extensive experiments on the FineDance dataset demonstrate state-of-the-art performance.
title MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
topic Sound
Graphics
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2505.14222