VTONQA: A Multi-Dimensional Quality Assessment Dataset for Virtual Try-on

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Xinyi, Wu, Sijing, Xu, Zitong, Li, Yunhao, Duan, Huiyu, Min, Xiongkuo, Zhai, Guangtao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918274073100288
author Wei, Xinyi
Wu, Sijing
Xu, Zitong
Li, Yunhao
Duan, Huiyu
Min, Xiongkuo
Zhai, Guangtao
author_facet Wei, Xinyi
Wu, Sijing
Xu, Zitong
Li, Yunhao
Duan, Huiyu
Min, Xiongkuo
Zhai, Guangtao
contents With the rapid development of e-commerce and digital fashion, image-based virtual try-on (VTON) has attracted increasing attention. However, existing VTON models often suffer from artifacts such as garment distortion and body inconsistency, highlighting the need for reliable quality evaluation of VTON-generated images. To this end, we construct VTONQA, the first multi-dimensional quality assessment dataset specifically designed for VTON, which contains 8,132 images generated by 11 representative VTON models, along with 24,396 mean opinion scores (MOSs) across three evaluation dimensions (i.e., clothing fit, body compatibility, and overall quality). Based on VTONQA, we benchmark both VTON models and a diverse set of image quality assessment (IQA) metrics, revealing the limitations of existing methods and highlighting the value of the proposed dataset. We believe that the VTONQA dataset and corresponding benchmarks will provide a solid foundation for perceptually aligned evaluation, benefiting both the development of quality assessment methods and the advancement of VTON models.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02945
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VTONQA: A Multi-Dimensional Quality Assessment Dataset for Virtual Try-on
Wei, Xinyi
Wu, Sijing
Xu, Zitong
Li, Yunhao
Duan, Huiyu
Min, Xiongkuo
Zhai, Guangtao
Computer Vision and Pattern Recognition
With the rapid development of e-commerce and digital fashion, image-based virtual try-on (VTON) has attracted increasing attention. However, existing VTON models often suffer from artifacts such as garment distortion and body inconsistency, highlighting the need for reliable quality evaluation of VTON-generated images. To this end, we construct VTONQA, the first multi-dimensional quality assessment dataset specifically designed for VTON, which contains 8,132 images generated by 11 representative VTON models, along with 24,396 mean opinion scores (MOSs) across three evaluation dimensions (i.e., clothing fit, body compatibility, and overall quality). Based on VTONQA, we benchmark both VTON models and a diverse set of image quality assessment (IQA) metrics, revealing the limitations of existing methods and highlighting the value of the proposed dataset. We believe that the VTONQA dataset and corresponding benchmarks will provide a solid foundation for perceptually aligned evaluation, benefiting both the development of quality assessment methods and the advancement of VTON models.
title VTONQA: A Multi-Dimensional Quality Assessment Dataset for Virtual Try-on
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.02945