Saved in:
Bibliographic Details
Main Authors: Zhang, Xulu, Wei, Xiao-Yong, Wu, Jinlin, Zhang, Tianyi, Zhang, Zhaoxiang, Lei, Zhen, Li, Qing
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2312.08048
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929206344024064
author Zhang, Xulu
Wei, Xiao-Yong
Wu, Jinlin
Zhang, Tianyi
Zhang, Zhaoxiang
Lei, Zhen
Li, Qing
author_facet Zhang, Xulu
Wei, Xiao-Yong
Wu, Jinlin
Zhang, Tianyi
Zhang, Zhaoxiang
Lei, Zhen
Li, Qing
contents Inversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of inverted concepts leads to the absence of other desired concepts. It stems from the fact that during inversion, the irrelevant semantics in the user images are also encoded, forcing the inverted concepts to occupy locations far from the core distribution in the embedding space. To address this issue, we propose a method that guides the inversion process towards the core distribution for compositional embeddings. Additionally, we introduce a spatial regularization approach to balance the attention on the concepts being composed. Our method is designed as a post-training approach and can be seamlessly integrated with other inversion methods. Experimental results demonstrate the effectiveness of our proposed approach in mitigating the overfitting problem and generating more diverse and balanced compositions of concepts in the synthesized images. The source code is available at https://github.com/zhangxulu1996/Compositional-Inversion.
format Preprint
id arxiv_https___arxiv_org_abs_2312_08048
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Compositional Inversion for Stable Diffusion Models
Zhang, Xulu
Wei, Xiao-Yong
Wu, Jinlin
Zhang, Tianyi
Zhang, Zhaoxiang
Lei, Zhen
Li, Qing
Computer Vision and Pattern Recognition
Inversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of inverted concepts leads to the absence of other desired concepts. It stems from the fact that during inversion, the irrelevant semantics in the user images are also encoded, forcing the inverted concepts to occupy locations far from the core distribution in the embedding space. To address this issue, we propose a method that guides the inversion process towards the core distribution for compositional embeddings. Additionally, we introduce a spatial regularization approach to balance the attention on the concepts being composed. Our method is designed as a post-training approach and can be seamlessly integrated with other inversion methods. Experimental results demonstrate the effectiveness of our proposed approach in mitigating the overfitting problem and generating more diverse and balanced compositions of concepts in the synthesized images. The source code is available at https://github.com/zhangxulu1996/Compositional-Inversion.
title Compositional Inversion for Stable Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.08048