AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niu, Yunfang, Wu, Lingxiang, Yi, Dong, Peng, Jie, Jiang, Ning, Wu, Haiying, Wang, Jinqiao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909352406810624
author Niu, Yunfang
Wu, Lingxiang
Yi, Dong
Peng, Jie
Jiang, Ning
Wu, Haiying
Wang, Jinqiao
author_facet Niu, Yunfang
Wu, Lingxiang
Yi, Dong
Peng, Jie
Jiang, Ning
Wu, Haiying
Wang, Jinqiao
contents Fashion image editing aims to modify a person's appearance based on a given instruction. Existing methods require auxiliary tools like segmenters and keypoint extractors, lacking a flexible and unified framework. Moreover, these methods are limited in the variety of clothing types they can handle, as most datasets focus on people in clean backgrounds and only include generic garments such as tops, pants, and dresses. These limitations restrict their applicability in real-world scenarios. In this paper, we first extend an existing dataset for human generation to include a wider range of apparel and more complex backgrounds. This extended dataset features people wearing diverse items such as tops, pants, dresses, skirts, headwear, scarves, shoes, socks, and bags. Additionally, we propose AnyDesign, a diffusion-based method that enables mask-free editing on versatile areas. Users can simply input a human image along with a corresponding prompt in either text or image format. Our approach incorporates Fashion DiT, equipped with a Fashion-Guidance Attention (FGA) module designed to fuse explicit apparel types and CLIP-encoded apparel features. Both Qualitative and quantitative experiments demonstrate that our method delivers high-quality fashion editing and outperforms contemporary text-guided fashion editing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion
Niu, Yunfang
Wu, Lingxiang
Yi, Dong
Peng, Jie
Jiang, Ning
Wu, Haiying
Wang, Jinqiao
Computer Vision and Pattern Recognition
Fashion image editing aims to modify a person's appearance based on a given instruction. Existing methods require auxiliary tools like segmenters and keypoint extractors, lacking a flexible and unified framework. Moreover, these methods are limited in the variety of clothing types they can handle, as most datasets focus on people in clean backgrounds and only include generic garments such as tops, pants, and dresses. These limitations restrict their applicability in real-world scenarios. In this paper, we first extend an existing dataset for human generation to include a wider range of apparel and more complex backgrounds. This extended dataset features people wearing diverse items such as tops, pants, dresses, skirts, headwear, scarves, shoes, socks, and bags. Additionally, we propose AnyDesign, a diffusion-based method that enables mask-free editing on versatile areas. Users can simply input a human image along with a corresponding prompt in either text or image format. Our approach incorporates Fashion DiT, equipped with a Fashion-Guidance Attention (FGA) module designed to fuse explicit apparel types and CLIP-encoded apparel features. Both Qualitative and quantitative experiments demonstrate that our method delivers high-quality fashion editing and outperforms contemporary text-guided fashion editing methods.
title AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.11553