OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kong, Zhe, Zhang, Yong, Yang, Tianyu, Wang, Tao, Zhang, Kaihao, Wu, Bizhu, Chen, Guanying, Liu, Wei, Luo, Wenhan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913438913003520
author Kong, Zhe
Zhang, Yong
Yang, Tianyu
Wang, Tao
Zhang, Kaihao
Wu, Bizhu
Chen, Guanying
Liu, Wei
Luo, Wenhan
author_facet Kong, Zhe
Zhang, Yong
Yang, Tianyu
Wang, Tao
Zhang, Kaihao
Wu, Bizhu
Chen, Guanying
Liu, Wei
Luo, Wenhan
contents Personalization is an important topic in text-to-image generation, especially the challenging multi-concept personalization. Current multi-concept methods are struggling with identity preservation, occlusion, and the harmony between foreground and background. In this work, we propose OMG, an occlusion-friendly personalized generation framework designed to seamlessly integrate multiple concepts within a single image. We propose a novel two-stage sampling solution. The first stage takes charge of layout generation and visual comprehension information collection for handling occlusions. The second one utilizes the acquired visual comprehension information and the designed noise blending to integrate multiple concepts while considering occlusions. We also observe that the initiation denoising timestep for noise blending is the key to identity preservation and layout. Moreover, our method can be combined with various single-concept models, such as LoRA and InstantID without additional tuning. Especially, LoRA models on civitai.com can be exploited directly. Extensive experiments demonstrate that OMG exhibits superior performance in multi-concept personalization.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10983
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models
Kong, Zhe
Zhang, Yong
Yang, Tianyu
Wang, Tao
Zhang, Kaihao
Wu, Bizhu
Chen, Guanying
Liu, Wei
Luo, Wenhan
Computer Vision and Pattern Recognition
Personalization is an important topic in text-to-image generation, especially the challenging multi-concept personalization. Current multi-concept methods are struggling with identity preservation, occlusion, and the harmony between foreground and background. In this work, we propose OMG, an occlusion-friendly personalized generation framework designed to seamlessly integrate multiple concepts within a single image. We propose a novel two-stage sampling solution. The first stage takes charge of layout generation and visual comprehension information collection for handling occlusions. The second one utilizes the acquired visual comprehension information and the designed noise blending to integrate multiple concepts while considering occlusions. We also observe that the initiation denoising timestep for noise blending is the key to identity preservation and layout. Moreover, our method can be combined with various single-concept models, such as LoRA and InstantID without additional tuning. Especially, LoRA models on civitai.com can be exploited directly. Extensive experiments demonstrate that OMG exhibits superior performance in multi-concept personalization.
title OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.10983