Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jun-Kun, Bansal, Aayush, Vo, Minh Phuoc, Wang, Yu-Xiong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912570471874560
author Chen, Jun-Kun
Bansal, Aayush
Vo, Minh Phuoc
Wang, Yu-Xiong
author_facet Chen, Jun-Kun
Bansal, Aayush
Vo, Minh Phuoc
Wang, Yu-Xiong
contents We introduce the Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an auto-regressive, segment-by-segment generation process, eliminating the need for resource-intensive generation and lengthy video data, while providing the flexibility to generate videos of arbitrary length. The key challenges of this task are twofold: ensuring local smoothness between adjacent segments and maintaining global temporal consistency across different segments. To address these challenges, we propose our VFR framework, which ensures smoothness through a prefix video condition and enforces consistency with the anchor video -- a 360-degree video that comprehensively captures the human's wholebody appearance. Our VFR generates minute-scale virtual try-on videos with both local smoothness and global temporal consistency under various motions, making it a pioneering work in long virtual try-on video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview
Chen, Jun-Kun
Bansal, Aayush
Vo, Minh Phuoc
Wang, Yu-Xiong
Computer Vision and Pattern Recognition
Machine Learning
We introduce the Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an auto-regressive, segment-by-segment generation process, eliminating the need for resource-intensive generation and lengthy video data, while providing the flexibility to generate videos of arbitrary length. The key challenges of this task are twofold: ensuring local smoothness between adjacent segments and maintaining global temporal consistency across different segments. To address these challenges, we propose our VFR framework, which ensures smoothness through a prefix video condition and enforces consistency with the anchor video -- a 360-degree video that comprehensively captures the human's wholebody appearance. Our VFR generates minute-scale virtual try-on videos with both local smoothness and global temporal consistency under various motions, making it a pioneering work in long virtual try-on video generation.
title Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.04450