Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeong, Dasol, Kang, Donggoo, Park, Jiwon, Lee, Hyebean, Paik, Joonki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909616732897280
author Jeong, Dasol
Kang, Donggoo
Park, Jiwon
Lee, Hyebean
Paik, Joonki
author_facet Jeong, Dasol
Kang, Donggoo
Park, Jiwon
Lee, Hyebean
Paik, Joonki
contents We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific null-text embeddings to preserve the structural integrity of the source image. By introducing a stage-wise latent injection strategy-shape injection in early steps and attribute injection in later steps-we enable precise, fine-grained modifications while maintaining global consistency. Cross-attention with reference latents facilitates semantic alignment between the source and reference. Extensive experiments across expression transfer, texture transformation, and style infusion demonstrate state-of-the-art performance, confirming the method's scalability and adaptability to diverse image editing scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models
Jeong, Dasol
Kang, Donggoo
Park, Jiwon
Lee, Hyebean
Paik, Joonki
Computer Vision and Pattern Recognition
We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific null-text embeddings to preserve the structural integrity of the source image. By introducing a stage-wise latent injection strategy-shape injection in early steps and attribute injection in later steps-we enable precise, fine-grained modifications while maintaining global consistency. Cross-attention with reference latents facilitates semantic alignment between the source and reference. Extensive experiments across expression transfer, texture transformation, and style infusion demonstrate state-of-the-art performance, confirming the method's scalability and adaptability to diverse image editing scenarios.
title Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.15723