DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junhao, Chen, Mingjin, Xu, Jianjin, Li, Xiang, Dong, Junting, Sun, Mingze, Jiang, Puhua, Li, Hongxiang, Yang, Yuhang, Zhao, Hao, Long, Xiaoxiao, Huang, Ruqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916755155189760
author Chen, Junhao
Chen, Mingjin
Xu, Jianjin
Li, Xiang
Dong, Junting
Sun, Mingze
Jiang, Puhua
Li, Hongxiang
Yang, Yuhang
Zhao, Hao
Long, Xiaoxiao
Huang, Ruqi
author_facet Chen, Junhao
Chen, Mingjin
Xu, Jianjin
Li, Xiang
Dong, Junting
Sun, Mingze
Jiang, Puhua
Li, Hongxiang
Yang, Yuhang
Zhao, Hao
Long, Xiaoxiao
Huang, Ruqi
contents Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image plus independent pose-mask streams into long, photorealistic videos while strictly preserving every identity. A novel MaskPoseAdapter binds "who" and "how" at every denoising step by fusing robust tracking masks with semantically rich-but noisy-pose heat-maps, eliminating the identity drift and appearance bleeding that plague frame-wise pipelines. To train and evaluate at scale, we introduce (i) PairFS-4K, 26 hours of dual-skater footage with 7,000+ distinct IDs, (ii) HumanRob-300, a one-hour humanoid-robot interaction set for rapid cross-domain transfer, and (iii) TogetherVideoBench, a three-track benchmark centered on the DanceTogEval-100 test suite covering dance, boxing, wrestling, yoga, and figure skating. On TogetherVideoBench, DanceTogether outperforms the prior arts by a significant margin. Moreover, we show that a one-hour fine-tune yields convincing human-robot videos, underscoring broad generalization to embodied-AI and HRI tasks. Extensive ablations confirm that persistent identity-action binding is critical to these gains. Together, our model, datasets, and benchmark lift CVG from single-subject choreography to compositionally controllable, multi-actor interaction, opening new avenues for digital production, simulation, and embodied intelligence. Our video demos and code are available at https://DanceTog.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
Chen, Junhao
Chen, Mingjin
Xu, Jianjin
Li, Xiang
Dong, Junting
Sun, Mingze
Jiang, Puhua
Li, Hongxiang
Yang, Yuhang
Zhao, Hao
Long, Xiaoxiao
Huang, Ruqi
Computer Vision and Pattern Recognition
Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image plus independent pose-mask streams into long, photorealistic videos while strictly preserving every identity. A novel MaskPoseAdapter binds "who" and "how" at every denoising step by fusing robust tracking masks with semantically rich-but noisy-pose heat-maps, eliminating the identity drift and appearance bleeding that plague frame-wise pipelines. To train and evaluate at scale, we introduce (i) PairFS-4K, 26 hours of dual-skater footage with 7,000+ distinct IDs, (ii) HumanRob-300, a one-hour humanoid-robot interaction set for rapid cross-domain transfer, and (iii) TogetherVideoBench, a three-track benchmark centered on the DanceTogEval-100 test suite covering dance, boxing, wrestling, yoga, and figure skating. On TogetherVideoBench, DanceTogether outperforms the prior arts by a significant margin. Moreover, we show that a one-hour fine-tune yields convincing human-robot videos, underscoring broad generalization to embodied-AI and HRI tasks. Extensive ablations confirm that persistent identity-action binding is critical to these gains. Together, our model, datasets, and benchmark lift CVG from single-subject choreography to compositionally controllable, multi-actor interaction, opening new avenues for digital production, simulation, and embodied intelligence. Our video demos and code are available at https://DanceTog.github.io/.
title DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18078