Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kwon, Gihyun, Jenni, Simon, Li, Dingzeyu, Lee, Joon-Young, Ye, Jong Chul, Heilbron, Fabian Caba
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909161909911552
author Kwon, Gihyun
Jenni, Simon
Li, Dingzeyu
Lee, Joon-Young
Ye, Jong Chul
Heilbron, Fabian Caba
author_facet Kwon, Gihyun
Jenni, Simon
Li, Dingzeyu
Lee, Joon-Young
Ye, Jong Chul
Heilbron, Fabian Caba
contents While there has been significant progress in customizing text-to-image generation models, generating images that combine multiple personalized concepts remains challenging. In this work, we introduce Concept Weaver, a method for composing customized text-to-image diffusion models at inference time. Specifically, the method breaks the process into two steps: creating a template image aligned with the semantics of input prompts, and then personalizing the template using a concept fusion strategy. The fusion strategy incorporates the appearance of the target concepts into the template image while retaining its structural details. The results indicate that our method can generate multiple custom concepts with higher identity fidelity compared to alternative approaches. Furthermore, the method is shown to seamlessly handle more than two concepts and closely follow the semantic meaning of the input prompt without blending appearances across different subjects.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models
Kwon, Gihyun
Jenni, Simon
Li, Dingzeyu
Lee, Joon-Young
Ye, Jong Chul
Heilbron, Fabian Caba
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
While there has been significant progress in customizing text-to-image generation models, generating images that combine multiple personalized concepts remains challenging. In this work, we introduce Concept Weaver, a method for composing customized text-to-image diffusion models at inference time. Specifically, the method breaks the process into two steps: creating a template image aligned with the semantics of input prompts, and then personalizing the template using a concept fusion strategy. The fusion strategy incorporates the appearance of the target concepts into the template image while retaining its structural details. The results indicate that our method can generate multiple custom concepts with higher identity fidelity compared to alternative approaches. Furthermore, the method is shown to seamlessly handle more than two concepts and closely follow the semantic meaning of the input prompt without blending appearances across different subjects.
title Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.03913