FashionComposer: Compositional Fashion Image Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ji, Sihui, Wang, Yiyang, Chen, Xi, Xu, Xiaogang, Luo, Hao, Zhao, Hengshuang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917873470930944
author Ji, Sihui
Wang, Yiyang
Chen, Xi
Xu, Xiaogang
Luo, Hao
Zhao, Hengshuang
author_facet Ji, Sihui
Wang, Yiyang
Chen, Xi
Xu, Xiaogang
Luo, Hao
Zhao, Hengshuang
contents We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and supports personalizing the appearance, pose, and figure of the human and assigning multiple garments in one pass. To achieve this, we first develop a universal framework capable of handling diverse input modalities. We construct scaled training data to enhance the model's robust compositional capabilities. To accommodate multiple reference images (garments and faces) seamlessly, we organize these references in a single image as an "asset library" and employ a reference UNet to extract appearance features. To inject the appearance features into the correct pixels in the generated result, we propose subject-binding attention. It binds the appearance features from different "assets" with the corresponding text features. In this way, the model could understand each asset according to their semantics, supporting arbitrary numbers and types of reference images. As a comprehensive solution, FashionComposer also supports many other applications like human album generation, diverse virtual try-on tasks, etc.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FashionComposer: Compositional Fashion Image Generation
Ji, Sihui
Wang, Yiyang
Chen, Xi
Xu, Xiaogang
Luo, Hao
Zhao, Hengshuang
Computer Vision and Pattern Recognition
We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and supports personalizing the appearance, pose, and figure of the human and assigning multiple garments in one pass. To achieve this, we first develop a universal framework capable of handling diverse input modalities. We construct scaled training data to enhance the model's robust compositional capabilities. To accommodate multiple reference images (garments and faces) seamlessly, we organize these references in a single image as an "asset library" and employ a reference UNet to extract appearance features. To inject the appearance features into the correct pixels in the generated result, we propose subject-binding attention. It binds the appearance features from different "assets" with the corresponding text features. In this way, the model could understand each asset according to their semantics, supporting arbitrary numbers and types of reference images. As a comprehensive solution, FashionComposer also supports many other applications like human album generation, diverse virtual try-on tasks, etc.
title FashionComposer: Compositional Fashion Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.14168