Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.21070 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908508756115456 |
|---|---|
| author | Chen, Jun-Kun Bansal, Aayush Vo, Minh Phuoc Wang, Yu-Xiong |
| author_facet | Chen, Jun-Kun Bansal, Aayush Vo, Minh Phuoc Wang, Yu-Xiong |
| contents | We present Dress&Dance, a video diffusion framework that generates high quality 5-second-long 24 FPS virtual try-on videos at 1152x720 resolution of a user wearing desired garments while moving in accordance with a given reference video. Our approach requires a single user image and supports a range of tops, bottoms, and one-piece garments, as well as simultaneous tops and bottoms try-on in a single pass. Key to our framework is CondNet, a novel conditioning network that leverages attention to unify multi-modal inputs (text, images, and videos), thereby enhancing garment registration and motion fidelity. CondNet is trained on heterogeneous training data, combining limited video data and a larger, more readily available image dataset, in a multistage progressive manner. Dress&Dance outperforms existing open source and commercial solutions and enables a high quality and flexible try-on experience. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_21070 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Dress&Dance: Dress up and Dance as You Like It - Technical Preview Chen, Jun-Kun Bansal, Aayush Vo, Minh Phuoc Wang, Yu-Xiong Computer Vision and Pattern Recognition Machine Learning We present Dress&Dance, a video diffusion framework that generates high quality 5-second-long 24 FPS virtual try-on videos at 1152x720 resolution of a user wearing desired garments while moving in accordance with a given reference video. Our approach requires a single user image and supports a range of tops, bottoms, and one-piece garments, as well as simultaneous tops and bottoms try-on in a single pass. Key to our framework is CondNet, a novel conditioning network that leverages attention to unify multi-modal inputs (text, images, and videos), thereby enhancing garment registration and motion fidelity. CondNet is trained on heterogeneous training data, combining limited video data and a larger, more readily available image dataset, in a multistage progressive manner. Dress&Dance outperforms existing open source and commercial solutions and enables a high quality and flexible try-on experience. |
| title | Dress&Dance: Dress up and Dance as You Like It - Technical Preview |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2508.21070 |