Saved in:
Bibliographic Details
Main Authors: Zhou, Yufan, Shen, Haoyu, Wang, Huan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.05606
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918191300608000
author Zhou, Yufan
Shen, Haoyu
Wang, Huan
author_facet Zhou, Yufan
Shen, Haoyu
Wang, Huan
contents Concept blending is a promising yet underexplored area in generative models. While recent approaches, such as embedding mixing and latent modification based on structural sketches, have been proposed, they often suffer from incompatible semantic information and discrepancies in shape and appearance. In this work, we introduce FreeBlend, an effective, training-free framework designed to address these challenges. To mitigate cross-modal loss and enhance feature detail, we leverage transferred image embeddings as conditional inputs. The framework employs a stepwise increasing interpolation strategy between latents, progressively adjusting the blending ratio to seamlessly integrate auxiliary features. Additionally, we introduce a feedback-driven mechanism that updates the auxiliary latents in reverse order, facilitating global blending and preventing rigid or unnatural outputs. Extensive experiments demonstrate that our method significantly improves both the semantic coherence and visual quality of blended images, yielding compelling and coherent results.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FreeBlend: Advancing Concept Blending with Staged Feedback-Driven Interpolation Diffusion
Zhou, Yufan
Shen, Haoyu
Wang, Huan
Computer Vision and Pattern Recognition
Concept blending is a promising yet underexplored area in generative models. While recent approaches, such as embedding mixing and latent modification based on structural sketches, have been proposed, they often suffer from incompatible semantic information and discrepancies in shape and appearance. In this work, we introduce FreeBlend, an effective, training-free framework designed to address these challenges. To mitigate cross-modal loss and enhance feature detail, we leverage transferred image embeddings as conditional inputs. The framework employs a stepwise increasing interpolation strategy between latents, progressively adjusting the blending ratio to seamlessly integrate auxiliary features. Additionally, we introduce a feedback-driven mechanism that updates the auxiliary latents in reverse order, facilitating global blending and preventing rigid or unnatural outputs. Extensive experiments demonstrate that our method significantly improves both the semantic coherence and visual quality of blended images, yielding compelling and coherent results.
title FreeBlend: Advancing Concept Blending with Staged Feedback-Driven Interpolation Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.05606