LoRACLR: Contrastive Adaptation for Customization of Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Simsar, Enis, Hofmann, Thomas, Tombari, Federico, Yanardag, Pinar
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908564261437440
author Simsar, Enis
Hofmann, Thomas
Tombari, Federico
Yanardag, Pinar
author_facet Simsar, Enis
Hofmann, Thomas
Tombari, Federico
Yanardag, Pinar
contents Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scenarios. However, current methods struggle with combining multiple personalized models, often leading to attribute entanglement or requiring separate training to preserve concept distinctiveness. We present LoRACLR, a novel approach for multi-concept image generation that merges multiple LoRA models, each fine-tuned for a distinct concept, into a single, unified model without additional individual fine-tuning. LoRACLR uses a contrastive objective to align and merge the weight spaces of these models, ensuring compatibility while minimizing interference. By enforcing distinct yet cohesive representations for each concept, LoRACLR enables efficient, scalable model composition for high-quality, multi-concept image synthesis. Our results highlight the effectiveness of LoRACLR in accurately merging multiple concepts, advancing the capabilities of personalized image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09622
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
Simsar, Enis
Hofmann, Thomas
Tombari, Federico
Yanardag, Pinar
Computer Vision and Pattern Recognition
Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scenarios. However, current methods struggle with combining multiple personalized models, often leading to attribute entanglement or requiring separate training to preserve concept distinctiveness. We present LoRACLR, a novel approach for multi-concept image generation that merges multiple LoRA models, each fine-tuned for a distinct concept, into a single, unified model without additional individual fine-tuning. LoRACLR uses a contrastive objective to align and merge the weight spaces of these models, ensuring compatibility while minimizing interference. By enforcing distinct yet cohesive representations for each concept, LoRACLR enables efficient, scalable model composition for high-quality, multi-concept image synthesis. Our results highlight the effectiveness of LoRACLR in accurately merging multiple concepts, advancing the capabilities of personalized image generation.
title LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09622