DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Sungnyun, Lee, Junsoo, Hong, Kibeom, Kim, Daesik, Ahn, Namhyuk
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911121156341760
author Kim, Sungnyun
Lee, Junsoo
Hong, Kibeom
Kim, Daesik
Ahn, Namhyuk
author_facet Kim, Sungnyun
Lee, Junsoo
Hong, Kibeom
Kim, Daesik
Ahn, Namhyuk
contents In this study, we aim to enhance the capabilities of diffusion-based text-to-image (T2I) generation models by integrating diverse modalities beyond textual descriptions within a unified framework. To this end, we categorize widely used conditional inputs into three modality types: structure, layout, and attribute. We propose a multimodal T2I diffusion model, which is capable of processing all three modalities within a single architecture without modifying the parameters of the pre-trained diffusion model, as only a small subset of components is updated. Our approach sets new benchmarks in multimodal generation through extensive quantitative and qualitative comparisons with existing conditional generation methods. We demonstrate that DiffBlender effectively integrates multiple sources of information and supports diverse applications in detailed image synthesis. The code and demo are available at https://github.com/sungnyun/diffblender.
format Preprint
id arxiv_https___arxiv_org_abs_2305_15194
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
Kim, Sungnyun
Lee, Junsoo
Hong, Kibeom
Kim, Daesik
Ahn, Namhyuk
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
In this study, we aim to enhance the capabilities of diffusion-based text-to-image (T2I) generation models by integrating diverse modalities beyond textual descriptions within a unified framework. To this end, we categorize widely used conditional inputs into three modality types: structure, layout, and attribute. We propose a multimodal T2I diffusion model, which is capable of processing all three modalities within a single architecture without modifying the parameters of the pre-trained diffusion model, as only a small subset of components is updated. Our approach sets new benchmarks in multimodal generation through extensive quantitative and qualitative comparisons with existing conditional generation methods. We demonstrate that DiffBlender effectively integrates multiple sources of information and supports diverse applications in detailed image synthesis. The code and demo are available at https://github.com/sungnyun/diffblender.
title DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2305.15194