Saved in:
Bibliographic Details
Main Authors: Chen, Jun-Kun, Bansal, Aayush, Vo, Minh Phuoc, Wang, Yu-Xiong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.21070
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908508756115456
author Chen, Jun-Kun
Bansal, Aayush
Vo, Minh Phuoc
Wang, Yu-Xiong
author_facet Chen, Jun-Kun
Bansal, Aayush
Vo, Minh Phuoc
Wang, Yu-Xiong
contents We present Dress&Dance, a video diffusion framework that generates high quality 5-second-long 24 FPS virtual try-on videos at 1152x720 resolution of a user wearing desired garments while moving in accordance with a given reference video. Our approach requires a single user image and supports a range of tops, bottoms, and one-piece garments, as well as simultaneous tops and bottoms try-on in a single pass. Key to our framework is CondNet, a novel conditioning network that leverages attention to unify multi-modal inputs (text, images, and videos), thereby enhancing garment registration and motion fidelity. CondNet is trained on heterogeneous training data, combining limited video data and a larger, more readily available image dataset, in a multistage progressive manner. Dress&Dance outperforms existing open source and commercial solutions and enables a high quality and flexible try-on experience.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dress&Dance: Dress up and Dance as You Like It - Technical Preview
Chen, Jun-Kun
Bansal, Aayush
Vo, Minh Phuoc
Wang, Yu-Xiong
Computer Vision and Pattern Recognition
Machine Learning
We present Dress&Dance, a video diffusion framework that generates high quality 5-second-long 24 FPS virtual try-on videos at 1152x720 resolution of a user wearing desired garments while moving in accordance with a given reference video. Our approach requires a single user image and supports a range of tops, bottoms, and one-piece garments, as well as simultaneous tops and bottoms try-on in a single pass. Key to our framework is CondNet, a novel conditioning network that leverages attention to unify multi-modal inputs (text, images, and videos), thereby enhancing garment registration and motion fidelity. CondNet is trained on heterogeneous training data, combining limited video data and a larger, more readily available image dataset, in a multistage progressive manner. Dress&Dance outperforms existing open source and commercial solutions and enables a high quality and flexible try-on experience.
title Dress&Dance: Dress up and Dance as You Like It - Technical Preview
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2508.21070