StyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoshikawa, Takato, Endo, Yuki, Kanamori, Yoshihiro
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913272615141376
author Yoshikawa, Takato
Endo, Yuki
Kanamori, Yoshihiro
author_facet Yoshikawa, Takato
Endo, Yuki
Kanamori, Yoshihiro
contents This paper tackles text-guided control of StyleGAN for editing garments in full-body human images. Existing StyleGAN-based methods suffer from handling the rich diversity of garments and body shapes and poses. We propose a framework for text-guided full-body human image synthesis via an attention-based latent code mapper, which enables more disentangled control of StyleGAN than existing mappers. Our latent code mapper adopts an attention mechanism that adaptively manipulates individual latent codes on different StyleGAN layers under text guidance. In addition, we introduce feature-space masking at inference time to avoid unwanted changes caused by text inputs. Our quantitative and qualitative evaluations reveal that our method can control generated images more faithfully to given texts than existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16759
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle StyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human
Yoshikawa, Takato
Endo, Yuki
Kanamori, Yoshihiro
Computer Vision and Pattern Recognition
Graphics
This paper tackles text-guided control of StyleGAN for editing garments in full-body human images. Existing StyleGAN-based methods suffer from handling the rich diversity of garments and body shapes and poses. We propose a framework for text-guided full-body human image synthesis via an attention-based latent code mapper, which enables more disentangled control of StyleGAN than existing mappers. Our latent code mapper adopts an attention mechanism that adaptively manipulates individual latent codes on different StyleGAN layers under text guidance. In addition, we introduce feature-space masking at inference time to avoid unwanted changes caused by text inputs. Our quantitative and qualitative evaluations reveal that our method can control generated images more faithfully to given texts than existing methods.
title StyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2305.16759