Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zhiqi, Xiong, Huixin, Wang, Haoyu, Wang, Longguang, Li, Zhiheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909163237408768
author Huang, Zhiqi
Xiong, Huixin
Wang, Haoyu
Wang, Longguang
Li, Zhiheng
author_facet Huang, Zhiqi
Xiong, Huixin
Wang, Haoyu
Wang, Longguang
Li, Zhiheng
contents Text-to-image generation has witnessed great progress, especially with the recent advancements in diffusion models. Since texts cannot provide detailed conditions like object appearance, reference images are usually leveraged for the control of objects in the generated images. However, existing methods still suffer limited accuracy when the relationship between the foreground and background is complicated. To address this issue, we develop a framework termed Mask-ControlNet by introducing an additional mask prompt. Specifically, we first employ large vision models to obtain masks to segment the objects of interest in the reference image. Then, the object images are employed as additional prompts to facilitate the diffusion model to better understand the relationship between foreground and background regions during image generation. Experiments show that the mask prompts enhance the controllability of the diffusion model to maintain higher fidelity to the reference image while achieving better image quality. Comparison with previous text-to-image generation methods demonstrates our method's superior quantitative and qualitative performance on the benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt
Huang, Zhiqi
Xiong, Huixin
Wang, Haoyu
Wang, Longguang
Li, Zhiheng
Computer Vision and Pattern Recognition
Text-to-image generation has witnessed great progress, especially with the recent advancements in diffusion models. Since texts cannot provide detailed conditions like object appearance, reference images are usually leveraged for the control of objects in the generated images. However, existing methods still suffer limited accuracy when the relationship between the foreground and background is complicated. To address this issue, we develop a framework termed Mask-ControlNet by introducing an additional mask prompt. Specifically, we first employ large vision models to obtain masks to segment the objects of interest in the reference image. Then, the object images are employed as additional prompts to facilitate the diffusion model to better understand the relationship between foreground and background regions during image generation. Experiments show that the mask prompts enhance the controllability of the diffusion model to maintain higher fidelity to the reference image while achieving better image quality. Comparison with previous text-to-image generation methods demonstrates our method's superior quantitative and qualitative performance on the benchmark datasets.
title Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.05331