IIDM: Image-to-Image Diffusion Model for Semantic Image Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Feng, Chang, Xiaobin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916362354425856
author Liu, Feng
Chang, Xiaobin
author_facet Liu, Feng
Chang, Xiaobin
contents Semantic image synthesis aims to generate high-quality images given semantic conditions, i.e. segmentation masks and style reference images. Existing methods widely adopt generative adversarial networks (GANs). GANs take all conditional inputs and directly synthesize images in a single forward step. In this paper, semantic image synthesis is treated as an image denoising task and is handled with a novel image-to-image diffusion model (IIDM). Specifically, the style reference is first contaminated with random noise and then progressively denoised by IIDM, guided by segmentation masks. Moreover, three techniques, refinement, color-transfer and model ensembles, are proposed to further boost the generation quality. They are plug-in inference modules and do not require additional training. Extensive experiments show that our IIDM outperforms existing state-of-the-art methods by clear margins. Further analysis is provided via detailed demonstrations. We have implemented IIDM based on the Jittor framework; code is available at https://github.com/ader47/jittor-jieke-semantic_images_synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13378
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IIDM: Image-to-Image Diffusion Model for Semantic Image Synthesis
Liu, Feng
Chang, Xiaobin
Computer Vision and Pattern Recognition
Semantic image synthesis aims to generate high-quality images given semantic conditions, i.e. segmentation masks and style reference images. Existing methods widely adopt generative adversarial networks (GANs). GANs take all conditional inputs and directly synthesize images in a single forward step. In this paper, semantic image synthesis is treated as an image denoising task and is handled with a novel image-to-image diffusion model (IIDM). Specifically, the style reference is first contaminated with random noise and then progressively denoised by IIDM, guided by segmentation masks. Moreover, three techniques, refinement, color-transfer and model ensembles, are proposed to further boost the generation quality. They are plug-in inference modules and do not require additional training. Extensive experiments show that our IIDM outperforms existing state-of-the-art methods by clear margins. Further analysis is provided via detailed demonstrations. We have implemented IIDM based on the Jittor framework; code is available at https://github.com/ader47/jittor-jieke-semantic_images_synthesis.
title IIDM: Image-to-Image Diffusion Model for Semantic Image Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.13378