SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Hung, Nguyen, Quang Qui-Vinh, Nguyen, Khoi, Nguyen, Rang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929637338120192
author Nguyen, Hung
Nguyen, Quang Qui-Vinh
Nguyen, Khoi
Nguyen, Rang
author_facet Nguyen, Hung
Nguyen, Quang Qui-Vinh
Nguyen, Khoi
Nguyen, Rang
contents Given an input video of a person and a new garment, the objective of this paper is to synthesize a new video where the person is wearing the specified garment while maintaining spatiotemporal consistency. Although significant advances have been made in image-based virtual try-on, extending these successes to video often leads to frame-to-frame inconsistencies. Some approaches have attempted to address this by increasing the overlap of frames across multiple video chunks, but this comes at a steep computational cost due to the repeated processing of the same frames, especially for long video sequences. To tackle these challenges, we reconceptualize video virtual try-on as a conditional video inpainting task, with garments serving as input conditions. Specifically, our approach enhances image diffusion models by incorporating temporal attention layers to improve temporal coherence. To reduce computational overhead, we propose ShiftCaching, a novel technique that maintains temporal consistency while minimizing redundant computations. Furthermore, we introduce the TikTokDress dataset, a new video try-on dataset featuring more complex backgrounds, challenging movements, and higher resolution compared to existing public datasets. Extensive experiments demonstrate that our approach outperforms current baselines, particularly in terms of video consistency and inference speed. The project page is available at https://swift-try.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
Nguyen, Hung
Nguyen, Quang Qui-Vinh
Nguyen, Khoi
Nguyen, Rang
Computer Vision and Pattern Recognition
Artificial Intelligence
Given an input video of a person and a new garment, the objective of this paper is to synthesize a new video where the person is wearing the specified garment while maintaining spatiotemporal consistency. Although significant advances have been made in image-based virtual try-on, extending these successes to video often leads to frame-to-frame inconsistencies. Some approaches have attempted to address this by increasing the overlap of frames across multiple video chunks, but this comes at a steep computational cost due to the repeated processing of the same frames, especially for long video sequences. To tackle these challenges, we reconceptualize video virtual try-on as a conditional video inpainting task, with garments serving as input conditions. Specifically, our approach enhances image diffusion models by incorporating temporal attention layers to improve temporal coherence. To reduce computational overhead, we propose ShiftCaching, a novel technique that maintains temporal consistency while minimizing redundant computations. Furthermore, we introduce the TikTokDress dataset, a new video try-on dataset featuring more complex backgrounds, challenging movements, and higher resolution compared to existing public datasets. Extensive experiments demonstrate that our approach outperforms current baselines, particularly in terms of video consistency and inference speed. The project page is available at https://swift-try.github.io/.
title SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.10178