WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929421245480960 |
|---|---|
| author | He, Zijian Chen, Peixin Wang, Guangrun Li, Guanbin Torr, Philip H. S. Lin, Liang |
| author_facet | He, Zijian Chen, Peixin Wang, Guangrun Li, Guanbin Torr, Philip H. S. Lin, Liang |
| contents | Video virtual try-on aims to generate realistic sequences that maintain garment identity and adapt to a person's pose and body shape in source videos. Traditional image-based methods, relying on warping and blending, struggle with complex human movements and occlusions, limiting their effectiveness in video try-on applications. Moreover, video-based models require extensive, high-quality data and substantial computational resources. To tackle these issues, we reconceptualize video try-on as a process of generating videos conditioned on garment descriptions and human motion. Our solution, WildVidFit, employs image-based controlled diffusion models for a streamlined, one-stage approach. This model, conditioned on specific garments and individuals, is trained on still images rather than videos. It leverages diffusion guidance from pre-trained models including a video masked autoencoder for segment smoothness improvement and a self-supervised model for feature alignment of adjacent frame in the latent space. This integration markedly boosts the model's ability to maintain temporal coherence, enabling more effective video try-on within an image-based framework. Our experiments on the VITON-HD and DressCode datasets, along with tests on the VVT and TikTok datasets, demonstrate WildVidFit's capability to generate fluid and coherent videos. The project page website is at wildvidfit-project.github.io. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_10625 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models He, Zijian Chen, Peixin Wang, Guangrun Li, Guanbin Torr, Philip H. S. Lin, Liang Computer Vision and Pattern Recognition Video virtual try-on aims to generate realistic sequences that maintain garment identity and adapt to a person's pose and body shape in source videos. Traditional image-based methods, relying on warping and blending, struggle with complex human movements and occlusions, limiting their effectiveness in video try-on applications. Moreover, video-based models require extensive, high-quality data and substantial computational resources. To tackle these issues, we reconceptualize video try-on as a process of generating videos conditioned on garment descriptions and human motion. Our solution, WildVidFit, employs image-based controlled diffusion models for a streamlined, one-stage approach. This model, conditioned on specific garments and individuals, is trained on still images rather than videos. It leverages diffusion guidance from pre-trained models including a video masked autoencoder for segment smoothness improvement and a self-supervised model for feature alignment of adjacent frame in the latent space. This integration markedly boosts the model's ability to maintain temporal coherence, enabling more effective video try-on within an image-based framework. Our experiments on the VITON-HD and DressCode datasets, along with tests on the VVT and TikTok datasets, demonstrate WildVidFit's capability to generate fluid and coherent videos. The project page website is at wildvidfit-project.github.io. |
| title | WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2407.10625 |