IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Xiaojing, Wu, Zhonghua, Yang, Xiaofeng, Lin, Guosheng, Wu, Qingyao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910801467539456
author Zhong, Xiaojing
Wu, Zhonghua
Yang, Xiaofeng
Lin, Guosheng
Wu, Qingyao
author_facet Zhong, Xiaojing
Wu, Zhonghua
Yang, Xiaofeng
Lin, Guosheng
Wu, Qingyao
contents Given a pair of images depicting a person and a garment separately, image-based 3D virtual try-on methods aim to reconstruct a 3D human model that realistically portrays the person wearing the desired garment. In this paper, we present IPVTON, a novel image-based 3D virtual try-on framework. IPVTON employs score distillation sampling with image prompts to optimize a hybrid 3D human representation, integrating target garment features into diffusion priors through an image prompt adapter. To avoid interference with non-target areas, we leverage mask-guided image prompt embeddings to focus the image features on the try-on regions. Moreover, we impose geometric constraints on the 3D model with a pseudo silhouette generated by ControlNet, ensuring that the clothed 3D human model retains the shape of the source identity while accurately wearing the target garments. Extensive qualitative and quantitative experiments demonstrate that IPVTON outperforms previous methods in image-based 3D virtual try-on tasks, excelling in both geometry and texture.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15616
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
Zhong, Xiaojing
Wu, Zhonghua
Yang, Xiaofeng
Lin, Guosheng
Wu, Qingyao
Computer Vision and Pattern Recognition
Given a pair of images depicting a person and a garment separately, image-based 3D virtual try-on methods aim to reconstruct a 3D human model that realistically portrays the person wearing the desired garment. In this paper, we present IPVTON, a novel image-based 3D virtual try-on framework. IPVTON employs score distillation sampling with image prompts to optimize a hybrid 3D human representation, integrating target garment features into diffusion priors through an image prompt adapter. To avoid interference with non-target areas, we leverage mask-guided image prompt embeddings to focus the image features on the try-on regions. Moreover, we impose geometric constraints on the 3D model with a pseudo silhouette generated by ControlNet, ensuring that the clothed 3D human model retains the shape of the source identity while accurately wearing the target garments. Extensive qualitative and quantitative experiments demonstrate that IPVTON outperforms previous methods in image-based 3D virtual try-on tasks, excelling in both geometry and texture.
title IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.15616