Fashionability-Enhancing Outfit Image Editing with Conditional Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Qice, Hirakawa, Yuki, Shimizu, Ryotaro, Furusawa, Takuya, Simo-Serra, Edgar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909440313131008
author Qin, Qice
Hirakawa, Yuki
Shimizu, Ryotaro
Furusawa, Takuya
Simo-Serra, Edgar
author_facet Qin, Qice
Hirakawa, Yuki
Shimizu, Ryotaro
Furusawa, Takuya
Simo-Serra, Edgar
contents Image generation in the fashion domain has predominantly focused on preserving body characteristics or following input prompts, but little attention has been paid to improving the inherent fashionability of the output images. This paper presents a novel diffusion model-based approach that generates fashion images with improved fashionability while maintaining control over key attributes. Key components of our method include: 1) fashionability enhancement, which ensures that the generated images are more fashionable than the input; 2) preservation of body characteristics, encouraging the generated images to maintain the original shape and proportions of the input; and 3) automatic fashion optimization, which does not rely on manual input or external prompts. We also employ two methods to collect training data for guidance while generating and evaluating the images. In particular, we rate outfit images using fashionability scores annotated by multiple fashion experts through OpenSkill-based and five critical aspect-based pairwise comparisons. These methods provide complementary perspectives for assessing and improving the fashionability of the generated images. The experimental results show that our approach outperforms the baseline Fashion++ in generating images with superior fashionability, demonstrating its effectiveness in producing more stylish and appealing fashion images.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18421
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fashionability-Enhancing Outfit Image Editing with Conditional Diffusion Models
Qin, Qice
Hirakawa, Yuki
Shimizu, Ryotaro
Furusawa, Takuya
Simo-Serra, Edgar
Computer Vision and Pattern Recognition
Image generation in the fashion domain has predominantly focused on preserving body characteristics or following input prompts, but little attention has been paid to improving the inherent fashionability of the output images. This paper presents a novel diffusion model-based approach that generates fashion images with improved fashionability while maintaining control over key attributes. Key components of our method include: 1) fashionability enhancement, which ensures that the generated images are more fashionable than the input; 2) preservation of body characteristics, encouraging the generated images to maintain the original shape and proportions of the input; and 3) automatic fashion optimization, which does not rely on manual input or external prompts. We also employ two methods to collect training data for guidance while generating and evaluating the images. In particular, we rate outfit images using fashionability scores annotated by multiple fashion experts through OpenSkill-based and five critical aspect-based pairwise comparisons. These methods provide complementary perspectives for assessing and improving the fashionability of the generated images. The experimental results show that our approach outperforms the baseline Fashion++ in generating images with superior fashionability, demonstrating its effectiveness in producing more stylish and appealing fashion images.
title Fashionability-Enhancing Outfit Image Editing with Conditional Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.18421