Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, Siqi, Chen, Jingwen, Pan, Yingwei, Yao, Ting, Mei, Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912387982950400
author Wan, Siqi
Chen, Jingwen
Pan, Yingwei
Yao, Ting
Mei, Tao
author_facet Wan, Siqi
Chen, Jingwen
Pan, Yingwei
Yao, Ting
Mei, Tao
contents Diffusion models have shown preliminary success in virtual try-on (VTON) task. The typical dual-branch architecture comprises two UNets for implicit garment deformation and synthesized image generation respectively, and has emerged as the recipe for VTON task. Nevertheless, the problem remains challenging to preserve the shape and every detail of the given garment due to the intrinsic stochasticity of diffusion model. To alleviate this issue, we novelly propose to explicitly capitalize on visual correspondence as the prior to tame diffusion process instead of simply feeding the whole garment into UNet as the appearance reference. Specifically, we interpret the fine-grained appearance and texture details as a set of structured semantic points, and match the semantic points rooted in garment to the ones over target person through local flow warping. Such 2D points are then augmented into 3D-aware cues with depth/normal map of target person. The correspondence mimics the way of putting clothing on human body and the 3D-aware cues act as semantic point matching to supervise diffusion model training. A point-focused diffusion loss is further devised to fully take the advantage of semantic point matching. Extensive experiments demonstrate strong garment detail preservation of our approach, evidenced by state-of-the-art VTON performances on both VITON-HD and DressCode datasets. Code is publicly available at: https://github.com/HiDream-ai/SPM-Diff.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16977
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
Wan, Siqi
Chen, Jingwen
Pan, Yingwei
Yao, Ting
Mei, Tao
Computer Vision and Pattern Recognition
Multimedia
Diffusion models have shown preliminary success in virtual try-on (VTON) task. The typical dual-branch architecture comprises two UNets for implicit garment deformation and synthesized image generation respectively, and has emerged as the recipe for VTON task. Nevertheless, the problem remains challenging to preserve the shape and every detail of the given garment due to the intrinsic stochasticity of diffusion model. To alleviate this issue, we novelly propose to explicitly capitalize on visual correspondence as the prior to tame diffusion process instead of simply feeding the whole garment into UNet as the appearance reference. Specifically, we interpret the fine-grained appearance and texture details as a set of structured semantic points, and match the semantic points rooted in garment to the ones over target person through local flow warping. Such 2D points are then augmented into 3D-aware cues with depth/normal map of target person. The correspondence mimics the way of putting clothing on human body and the 3D-aware cues act as semantic point matching to supervise diffusion model training. A point-focused diffusion loss is further devised to fully take the advantage of semantic point matching. Extensive experiments demonstrate strong garment detail preservation of our approach, evidenced by state-of-the-art VTON performances on both VITON-HD and DressCode datasets. Code is publicly available at: https://github.com/HiDream-ai/SPM-Diff.
title Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2505.16977