OmniVTON: Training-Free Universal Virtual Try-On

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhaotong, Li, Yuhui, He, Shengfeng, Li, Xinzhe, Xu, Yangyang, Dong, Junyu, Du, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915401243295744
author Yang, Zhaotong
Li, Yuhui
He, Shengfeng
Li, Xinzhe
Xu, Yangyang
Dong, Junyu
Du, Yong
author_facet Yang, Zhaotong
Li, Yuhui
He, Shengfeng
Li, Xinzhe
Xu, Yangyang
Dong, Junyu
Du, Yong
contents Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve fine-grained texture retention. For precise pose alignment, we utilize DDIM inversion to capture structural cues while suppressing texture interference, ensuring accurate body alignment independent of the original image textures. By disentangling garment and pose constraints, OmniVTON eliminates the bias inherent in diffusion models when handling multiple conditions simultaneously. Experimental results demonstrate that OmniVTON achieves superior performance across diverse datasets, garment types, and application scenarios. Notably, it is the first framework capable of multi-human VTON, enabling realistic garment transfer across multiple individuals in a single scene. Code is available at https://github.com/Jerome-Young/OmniVTON
format Preprint
id arxiv_https___arxiv_org_abs_2507_15037
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniVTON: Training-Free Universal Virtual Try-On
Yang, Zhaotong
Li, Yuhui
He, Shengfeng
Li, Xinzhe
Xu, Yangyang
Dong, Junyu
Du, Yong
Computer Vision and Pattern Recognition
Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve fine-grained texture retention. For precise pose alignment, we utilize DDIM inversion to capture structural cues while suppressing texture interference, ensuring accurate body alignment independent of the original image textures. By disentangling garment and pose constraints, OmniVTON eliminates the bias inherent in diffusion models when handling multiple conditions simultaneously. Experimental results demonstrate that OmniVTON achieves superior performance across diverse datasets, garment types, and application scenarios. Notably, it is the first framework capable of multi-human VTON, enabling realistic garment transfer across multiple individuals in a single scene. Code is available at https://github.com/Jerome-Young/OmniVTON
title OmniVTON: Training-Free Universal Virtual Try-On
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.15037