Video Virtual Try-on with Conditional Diffusion Transformer Inpainter

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zou, Cheng, Cheng, Senlin, Xu, Bolei, Zheng, Dandan, Li, Xiaobo, Chen, Jingdong, Yang, Ming
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911024378019840
author Zou, Cheng
Cheng, Senlin
Xu, Bolei
Zheng, Dandan
Li, Xiaobo
Chen, Jingdong
Yang, Ming
author_facet Zou, Cheng
Cheng, Senlin
Xu, Bolei
Zheng, Dandan
Li, Xiaobo
Chen, Jingdong
Yang, Ming
contents Video virtual try-on aims to naturally fit a garment to a target person in consecutive video frames. It is a challenging task, on the one hand, the output video should be in good spatial-temporal consistency, on the other hand, the details of the given garment need to be preserved well in all the frames. Naively using image-based try-on methods frame by frame can get poor results due to severe inconsistency. Recent diffusion-based video try-on methods, though very few, happen to coincide with a similar solution: inserting temporal attention into image-based try-on model to adapt it for video try-on task, which have shown improvements but there still exist inconsistency problems. In this paper, we propose ViTI (Video Try-on Inpainter), formulate and implement video virtual try-on as a conditional video inpainting task, which is different from previous methods. In this way, we start with a video generation problem instead of an image-based try-on problem, which from the beginning has a better spatial-temporal consistency. Specifically, at first we build a video inpainting framework based on Diffusion Transformer with full 3D spatial-temporal attention, and then we progressively adapt it for video garment inpainting, with a collection of masking strategies and multi-stage training. After these steps, the model can inpaint the masked garment area with appropriate garment pixels according to the prompt with good spatial-temporal consistency. Finally, as other try-on methods, garment condition is added to the model to make sure the inpainted garment appearance and details are as expected. Both quantitative and qualitative experimental results show that ViTI is superior to previous works.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
Zou, Cheng
Cheng, Senlin
Xu, Bolei
Zheng, Dandan
Li, Xiaobo
Chen, Jingdong
Yang, Ming
Computer Vision and Pattern Recognition
Video virtual try-on aims to naturally fit a garment to a target person in consecutive video frames. It is a challenging task, on the one hand, the output video should be in good spatial-temporal consistency, on the other hand, the details of the given garment need to be preserved well in all the frames. Naively using image-based try-on methods frame by frame can get poor results due to severe inconsistency. Recent diffusion-based video try-on methods, though very few, happen to coincide with a similar solution: inserting temporal attention into image-based try-on model to adapt it for video try-on task, which have shown improvements but there still exist inconsistency problems. In this paper, we propose ViTI (Video Try-on Inpainter), formulate and implement video virtual try-on as a conditional video inpainting task, which is different from previous methods. In this way, we start with a video generation problem instead of an image-based try-on problem, which from the beginning has a better spatial-temporal consistency. Specifically, at first we build a video inpainting framework based on Diffusion Transformer with full 3D spatial-temporal attention, and then we progressively adapt it for video garment inpainting, with a collection of masking strategies and multi-stage training. After these steps, the model can inpaint the masked garment area with appropriate garment pixels according to the prompt with good spatial-temporal consistency. Finally, as other try-on methods, garment condition is added to the model to make sure the inpainted garment appearance and details are as expected. Both quantitative and qualitative experimental results show that ViTI is superior to previous works.
title Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21270