Towards Generalized and Training-Free Text-Guided Semantic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Yu, Cai, Xiao, Zeng, Pengpeng, Zhang, Shuai, Song, Jingkuan, Gao, Lianli, Shen, Heng Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916819117277184
author Hong, Yu
Cai, Xiao
Zeng, Pengpeng
Zhang, Shuai
Song, Jingkuan
Gao, Lianli
Shen, Heng Tao
author_facet Hong, Yu
Cai, Xiao
Zeng, Pengpeng
Zhang, Shuai
Song, Jingkuan
Gao, Lianli
Shen, Heng Tao
contents Text-guided semantic manipulation refers to semantically editing an image generated from a source prompt to match a target prompt, enabling the desired semantic changes (e.g., addition, removal, and style transfer) while preserving irrelevant contents. With the powerful generative capabilities of the diffusion model, the task has shown the potential to generate high-fidelity visual content. Nevertheless, existing methods either typically require time-consuming fine-tuning (inefficient), fail to accomplish multiple semantic manipulations (poorly extensible), and/or lack support for different modality tasks (limited generalizability). Upon further investigation, we find that the geometric properties of noises in the diffusion model are strongly correlated with the semantic changes. Motivated by this, we propose a novel $\textit{GTF}$ for text-guided semantic manipulation, which has the following attractive capabilities: 1) $\textbf{Generalized}$: our $\textit{GTF}$ supports multiple semantic manipulations (e.g., addition, removal, and style transfer) and can be seamlessly integrated into all diffusion-based methods (i.e., Plug-and-play) across different modalities (i.e., modality-agnostic); and 2) $\textbf{Training-free}$: $\textit{GTF}$ produces high-fidelity results via simply controlling the geometric relationship between noises without tuning or optimization. Our extensive experiments demonstrate the efficacy of our approach, highlighting its potential to advance the state-of-the-art in semantics manipulation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Generalized and Training-Free Text-Guided Semantic Manipulation
Hong, Yu
Cai, Xiao
Zeng, Pengpeng
Zhang, Shuai
Song, Jingkuan
Gao, Lianli
Shen, Heng Tao
Computer Vision and Pattern Recognition
Text-guided semantic manipulation refers to semantically editing an image generated from a source prompt to match a target prompt, enabling the desired semantic changes (e.g., addition, removal, and style transfer) while preserving irrelevant contents. With the powerful generative capabilities of the diffusion model, the task has shown the potential to generate high-fidelity visual content. Nevertheless, existing methods either typically require time-consuming fine-tuning (inefficient), fail to accomplish multiple semantic manipulations (poorly extensible), and/or lack support for different modality tasks (limited generalizability). Upon further investigation, we find that the geometric properties of noises in the diffusion model are strongly correlated with the semantic changes. Motivated by this, we propose a novel $\textit{GTF}$ for text-guided semantic manipulation, which has the following attractive capabilities: 1) $\textbf{Generalized}$: our $\textit{GTF}$ supports multiple semantic manipulations (e.g., addition, removal, and style transfer) and can be seamlessly integrated into all diffusion-based methods (i.e., Plug-and-play) across different modalities (i.e., modality-agnostic); and 2) $\textbf{Training-free}$: $\textit{GTF}$ produces high-fidelity results via simply controlling the geometric relationship between noises without tuning or optimization. Our extensive experiments demonstrate the efficacy of our approach, highlighting its potential to advance the state-of-the-art in semantics manipulation.
title Towards Generalized and Training-Free Text-Guided Semantic Manipulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.17269