CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsin-Ying, Lee, Chan, Kelvin C. K., Yang, Ming-Hsuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915652440162304
author Hsin-Ying, Lee
Chan, Kelvin C. K.
Yang, Ming-Hsuan
author_facet Hsin-Ying, Lee
Chan, Kelvin C. K.
Yang, Ming-Hsuan
contents While text-to-image generative models can synthesize diverse and faithful content, subject variation across multiple generations limits their application to long-form content generation. Existing approaches require time-consuming fine-tuning, reference images for all subjects, or access to previously generated content. We introduce Contrastive Concept Instantiation (CoCoIns), a framework that effectively synthesizes consistent subjects across multiple independent generations. The framework consists of a generative model and a mapping network that transforms input latent codes into pseudo-words associated with specific concept instances. Users can generate consistent subjects by reusing the same latent codes. To construct such associations, we propose a contrastive learning approach that trains the network to distinguish between different combinations of prompts and latent codes. Extensive evaluations on human faces with a single subject show that CoCoIns performs comparably to existing methods while maintaining greater flexibility. We also demonstrate the potential for extending CoCoIns to multiple subjects and other object categories.
format Preprint
id arxiv_https___arxiv_org_abs_2503_24387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
Hsin-Ying, Lee
Chan, Kelvin C. K.
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
While text-to-image generative models can synthesize diverse and faithful content, subject variation across multiple generations limits their application to long-form content generation. Existing approaches require time-consuming fine-tuning, reference images for all subjects, or access to previously generated content. We introduce Contrastive Concept Instantiation (CoCoIns), a framework that effectively synthesizes consistent subjects across multiple independent generations. The framework consists of a generative model and a mapping network that transforms input latent codes into pseudo-words associated with specific concept instances. Users can generate consistent subjects by reusing the same latent codes. To construct such associations, we propose a contrastive learning approach that trains the network to distinguish between different combinations of prompts and latent codes. Extensive evaluations on human faces with a single subject show that CoCoIns performs comparably to existing methods while maintaining greater flexibility. We also demonstrate the potential for extending CoCoIns to multiple subjects and other object categories.
title CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.24387