Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Minseo, Kwon, Minchan, Lee, Dongyeun, Jeon, Yunho, Kim, Junmo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908484359946240
author Kim, Minseo
Kwon, Minchan
Lee, Dongyeun
Jeon, Yunho
Kim, Junmo
author_facet Kim, Minseo
Kwon, Minchan
Lee, Dongyeun
Jeon, Yunho
Kim, Junmo
contents The recent demand for customized image generation raises a need for techniques that effectively extract the common concept from small sets of images. Existing methods typically rely on additional guidance, such as text prompts or spatial masks, to capture the common target concept. Unfortunately, relying on manually provided guidance can lead to incomplete separation of auxiliary features, which degrades generation quality.In this paper, we propose Contrastive Inversion, a novel approach that identifies the common concept by comparing the input images without relying on additional information. We train the target token along with the image-wise auxiliary text tokens via contrastive learning, which extracts the well-disentangled true semantics of the target. Then we apply disentangled cross-attention fine-tuning to improve concept fidelity without overfitting. Experimental results and analysis demonstrate that our method achieves a balanced, high-level performance in both concept representation and editing, outperforming existing techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07755
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion
Kim, Minseo
Kwon, Minchan
Lee, Dongyeun
Jeon, Yunho
Kim, Junmo
Computer Vision and Pattern Recognition
The recent demand for customized image generation raises a need for techniques that effectively extract the common concept from small sets of images. Existing methods typically rely on additional guidance, such as text prompts or spatial masks, to capture the common target concept. Unfortunately, relying on manually provided guidance can lead to incomplete separation of auxiliary features, which degrades generation quality.In this paper, we propose Contrastive Inversion, a novel approach that identifies the common concept by comparing the input images without relying on additional information. We train the target token along with the image-wise auxiliary text tokens via contrastive learning, which extracts the well-disentangled true semantics of the target. Then we apply disentangled cross-attention fine-tuning to improve concept fidelity without overfitting. Experimental results and analysis demonstrate that our method achieves a balanced, high-level performance in both concept representation and editing, outperforming existing techniques.
title Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.07755