MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Prerit, Fotso-Puepi, Jason Alexander, Li, Zhengyuan, Mehta, Jay, Bera, Aniket
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909750767124480
author Gupta, Prerit
Fotso-Puepi, Jason Alexander
Li, Zhengyuan
Mehta, Jay
Bera, Aniket
author_facet Gupta, Prerit
Fotso-Puepi, Jason Alexander
Li, Zhengyuan
Mehta, Jay
Bera, Aniket
contents We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data performed by professional dancers, synchronized with music, and detailed with over 10K fine-grained natural language descriptions. The annotations capture a rich movement vocabulary, detailing spatial relationships, body movements, and rhythm, making MDD the first dataset to seamlessly integrate human motions, music, and text for duet dance generation. We introduce two novel tasks supported by our dataset: (1) Text-to-Duet, where given music and a textual prompt, both the leader and follower dance motion are generated (2) Text-to-Dance Accompaniment, where given music, textual prompt, and the leader's motion, the follower's motion is generated in a cohesive, text-aligned manner. We include baseline evaluations on both tasks to support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16911
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
Gupta, Prerit
Fotso-Puepi, Jason Alexander
Li, Zhengyuan
Mehta, Jay
Bera, Aniket
Graphics
Computer Vision and Pattern Recognition
Multimedia
Sound
We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data performed by professional dancers, synchronized with music, and detailed with over 10K fine-grained natural language descriptions. The annotations capture a rich movement vocabulary, detailing spatial relationships, body movements, and rhythm, making MDD the first dataset to seamlessly integrate human motions, music, and text for duet dance generation. We introduce two novel tasks supported by our dataset: (1) Text-to-Duet, where given music and a textual prompt, both the leader and follower dance motion are generated (2) Text-to-Dance Accompaniment, where given music, textual prompt, and the leader's motion, the follower's motion is generated in a cohesive, text-aligned manner. We include baseline evaluations on both tasks to support future research.
title MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
topic Graphics
Computer Vision and Pattern Recognition
Multimedia
Sound
url https://arxiv.org/abs/2508.16911