Saved in:
Bibliographic Details
Main Authors: Li, Yangyang, Liu, Daqing, Liu, Wu, He, Allen, Liu, Xinchen, Zhang, Yongdong, Jin, Guoqing
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2412.12242
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911578900660224
author Li, Yangyang
Liu, Daqing
Liu, Wu
He, Allen
Liu, Xinchen
Zhang, Yongdong
Jin, Guoqing
author_facet Li, Yangyang
Liu, Daqing
Liu, Wu
He, Allen
Liu, Xinchen
Zhang, Yongdong
Jin, Guoqing
contents Creative visual concept generation often draws inspiration from specific concepts in a reference image to produce relevant outcomes. However, existing methods are typically constrained to single-aspect concept generation or are easily disrupted by irrelevant concepts in multi-aspect concept scenarios, leading to concept confusion and hindering creative generation. To address this, we propose OmniPrism, a visual concept disentangling approach for creative image generation. Our method learns disentangled concept representations guided by natural language and trains a diffusion model to incorporate these concepts. We utilize the rich semantic space of a multimodal extractor to achieve concept disentanglement from given images and concept guidance. To disentangle concepts with different semantics, we construct a paired concept disentangled dataset (PCD-200K), where each pair shares the same concept such as content, style, and composition. We learn disentangled concept representations through our contrastive orthogonal disentangled (COD) training pipeline, which are then injected into additional diffusion cross-attention layers for generation. A set of block embeddings is designed to adapt each block's concept domain in the diffusion models. Extensive experiments demonstrate that our method can generate high-quality, concept-disentangled results with high fidelity to text prompts and desired concepts.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12242
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OmniPrism: Learning Disentangled Visual Concept for Image Generation
Li, Yangyang
Liu, Daqing
Liu, Wu
He, Allen
Liu, Xinchen
Zhang, Yongdong
Jin, Guoqing
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Creative visual concept generation often draws inspiration from specific concepts in a reference image to produce relevant outcomes. However, existing methods are typically constrained to single-aspect concept generation or are easily disrupted by irrelevant concepts in multi-aspect concept scenarios, leading to concept confusion and hindering creative generation. To address this, we propose OmniPrism, a visual concept disentangling approach for creative image generation. Our method learns disentangled concept representations guided by natural language and trains a diffusion model to incorporate these concepts. We utilize the rich semantic space of a multimodal extractor to achieve concept disentanglement from given images and concept guidance. To disentangle concepts with different semantics, we construct a paired concept disentangled dataset (PCD-200K), where each pair shares the same concept such as content, style, and composition. We learn disentangled concept representations through our contrastive orthogonal disentangled (COD) training pipeline, which are then injected into additional diffusion cross-attention layers for generation. A set of block embeddings is designed to adapt each block's concept domain in the diffusion models. Extensive experiments demonstrate that our method can generate high-quality, concept-disentangled results with high fidelity to text prompts and desired concepts.
title OmniPrism: Learning Disentangled Visual Concept for Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.12242