FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tran, Gia-Nghia, Che, Quang-Huy, Vu, Trong-Tai Dam, Pham, Bich-Nga, Nguyen, Vinh-Tiep, Le, Trung-Nghia, Tran, Minh-Triet
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909565025517568
author Tran, Gia-Nghia
Che, Quang-Huy
Vu, Trong-Tai Dam
Pham, Bich-Nga
Nguyen, Vinh-Tiep
Le, Trung-Nghia
Tran, Minh-Triet
author_facet Tran, Gia-Nghia
Che, Quang-Huy
Vu, Trong-Tai Dam
Pham, Bich-Nga
Nguyen, Vinh-Tiep
Le, Trung-Nghia
Tran, Minh-Triet
contents Generating multiple new concepts remains a challenging problem in the text-to-image task. Current methods often overfit when trained on a small number of samples and struggle with attribute leakage, particularly for class-similar subjects (e.g., two specific dogs). In this paper, we introduce Fuse-and-Refine (FaR), a novel approach that tackles these challenges through two key contributions: Concept Fusion technique and Localized Refinement loss function. Concept Fusion systematically augments the training data by separating reference subjects from backgrounds and recombining them into composite images to increase diversity. This augmentation technique tackles the overfitting problem by mitigating the narrow distribution of the limited training samples. In addition, Localized Refinement loss function is introduced to preserve subject representative attributes by aligning each concept's attention map to its correct region. This approach effectively prevents attribute leakage by ensuring that the diffusion model distinguishes similar subjects without mixing their attention maps during the denoising process. By fine-tuning specific modules at the same time, FaR balances the learning of new concepts with the retention of previously learned knowledge. Empirical results show that FaR not only prevents overfitting and attribute leakage while maintaining photorealism, but also outperforms other state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03292
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
Tran, Gia-Nghia
Che, Quang-Huy
Vu, Trong-Tai Dam
Pham, Bich-Nga
Nguyen, Vinh-Tiep
Le, Trung-Nghia
Tran, Minh-Triet
Computer Vision and Pattern Recognition
Generating multiple new concepts remains a challenging problem in the text-to-image task. Current methods often overfit when trained on a small number of samples and struggle with attribute leakage, particularly for class-similar subjects (e.g., two specific dogs). In this paper, we introduce Fuse-and-Refine (FaR), a novel approach that tackles these challenges through two key contributions: Concept Fusion technique and Localized Refinement loss function. Concept Fusion systematically augments the training data by separating reference subjects from backgrounds and recombining them into composite images to increase diversity. This augmentation technique tackles the overfitting problem by mitigating the narrow distribution of the limited training samples. In addition, Localized Refinement loss function is introduced to preserve subject representative attributes by aligning each concept's attention map to its correct region. This approach effectively prevents attribute leakage by ensuring that the diffusion model distinguishes similar subjects without mixing their attention maps during the denoising process. By fine-tuning specific modules at the same time, FaR balances the learning of new concepts with the retention of previously learned knowledge. Empirical results show that FaR not only prevents overfitting and attribute leakage while maintaining photorealism, but also outperforms other state-of-the-art methods.
title FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.03292