Semantic Image Synthesis with Unconditional Generator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chae, Jungwoo, Cho, Hyunin, Go, Sooyeon, Choi, Kyungmook, Uh, Youngjung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909116773957632
author Chae, Jungwoo
Cho, Hyunin
Go, Sooyeon
Choi, Kyungmook
Uh, Youngjung
author_facet Chae, Jungwoo
Cho, Hyunin
Go, Sooyeon
Choi, Kyungmook
Uh, Youngjung
contents Semantic image synthesis (SIS) aims to generate realistic images that match given semantic masks. Despite recent advances allowing high-quality results and precise spatial control, they require a massive semantic segmentation dataset for training the models. Instead, we propose to employ a pre-trained unconditional generator and rearrange its feature maps according to proxy masks. The proxy masks are prepared from the feature maps of random samples in the generator by simple clustering. The feature rearranger learns to rearrange original feature maps to match the shape of the proxy masks that are either from the original sample itself or from random samples. Then we introduce a semantic mapper that produces the proxy masks from various input conditions including semantic masks. Our method is versatile across various applications such as free-form spatial editing of real images, sketch-to-photo, and even scribble-to-photo. Experiments validate advantages of our method on a range of datasets: human faces, animal faces, and buildings.
format Preprint
id arxiv_https___arxiv_org_abs_2402_14395
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantic Image Synthesis with Unconditional Generator
Chae, Jungwoo
Cho, Hyunin
Go, Sooyeon
Choi, Kyungmook
Uh, Youngjung
Computer Vision and Pattern Recognition
Semantic image synthesis (SIS) aims to generate realistic images that match given semantic masks. Despite recent advances allowing high-quality results and precise spatial control, they require a massive semantic segmentation dataset for training the models. Instead, we propose to employ a pre-trained unconditional generator and rearrange its feature maps according to proxy masks. The proxy masks are prepared from the feature maps of random samples in the generator by simple clustering. The feature rearranger learns to rearrange original feature maps to match the shape of the proxy masks that are either from the original sample itself or from random samples. Then we introduce a semantic mapper that produces the proxy masks from various input conditions including semantic masks. Our method is versatile across various applications such as free-form spatial editing of real images, sketch-to-photo, and even scribble-to-photo. Experiments validate advantages of our method on a range of datasets: human faces, animal faces, and buildings.
title Semantic Image Synthesis with Unconditional Generator
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.14395