Semantic Granularity Navigation in Image Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Liangsi, Guo, Minzhe, Chen, Xuhang, Shi, Yang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918519862460416
author Lu, Liangsi
Guo, Minzhe
Chen, Xuhang
Shi, Yang
author_facet Lu, Liangsi
Guo, Minzhe
Chen, Xuhang
Shi, Yang
contents Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic editability and structural fidelity. We trace a primary cause of this limitation to the implicit coupling of edit progress with model scale in existing paradigms. Under this coupling, stronger edits typically require visiting noisier states, which spends computation on destabilizing layout before the semantic change is well localized. We introduce NaviEdit, a training-free inference-time controller that decouples edit progress from model scale traversal through a strict self-consistency contract. NaviEdit operates at the rollout level and leaves the underlying pretrained model unchanged. It treats scale as a control input and reallocates a fixed step budget toward semantically responsive intermediate scales instead of destructive high-noise regimes. Experiments show positive average gains across compatible editors and flow backbones, supporting decoupling as a portable inference-time control principle.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21190
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Semantic Granularity Navigation in Image Editing
Lu, Liangsi
Guo, Minzhe
Chen, Xuhang
Shi, Yang
Computer Vision and Pattern Recognition
Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic editability and structural fidelity. We trace a primary cause of this limitation to the implicit coupling of edit progress with model scale in existing paradigms. Under this coupling, stronger edits typically require visiting noisier states, which spends computation on destabilizing layout before the semantic change is well localized. We introduce NaviEdit, a training-free inference-time controller that decouples edit progress from model scale traversal through a strict self-consistency contract. NaviEdit operates at the rollout level and leaves the underlying pretrained model unchanged. It treats scale as a control input and reallocates a fixed step budget toward semantically responsive intermediate scales instead of destructive high-noise regimes. Experiments show positive average gains across compatible editors and flow backbones, supporting decoupling as a portable inference-time control principle.
title Semantic Granularity Navigation in Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.21190