CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Zhe, Chen, Honghua, Li, Peng, Wei, Mingqiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908840171143168
author Zhu, Zhe
Chen, Honghua
Li, Peng
Wei, Mingqiang
author_facet Zhu, Zhe
Chen, Honghua
Li, Peng
Wei, Mingqiang
contents Text-driven 3D editing seeks to modify 3D scenes according to textual descriptions, and most existing approaches tackle this by adapting pre-trained 2D image editors to multi-view inputs. However, without explicit control over multi-view information exchange, they often fail to maintain cross-view consistency, leading to insufficient edits and blurry details. We introduce CoreEditor, a novel framework for consistent text-to-3D editing. The key innovation is a correspondence-constrained attention mechanism that enforces precise interactions between pixels expected to remain consistent throughout the diffusion denoising process. Beyond relying solely on geometric alignment, we further incorporate semantic similarity estimated during denoising, enabling more reliable correspondence modeling and robust multi-view editing. In addition, we design a selective editing pipeline that allows users to choose preferred results from multiple candidates, offering greater flexibility and user control. Extensive experiments show that CoreEditor produces high-quality, 3D-consistent edits with sharper details, significantly outperforming prior methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
Zhu, Zhe
Chen, Honghua
Li, Peng
Wei, Mingqiang
Computer Vision and Pattern Recognition
Text-driven 3D editing seeks to modify 3D scenes according to textual descriptions, and most existing approaches tackle this by adapting pre-trained 2D image editors to multi-view inputs. However, without explicit control over multi-view information exchange, they often fail to maintain cross-view consistency, leading to insufficient edits and blurry details. We introduce CoreEditor, a novel framework for consistent text-to-3D editing. The key innovation is a correspondence-constrained attention mechanism that enforces precise interactions between pixels expected to remain consistent throughout the diffusion denoising process. Beyond relying solely on geometric alignment, we further incorporate semantic similarity estimated during denoising, enabling more reliable correspondence modeling and robust multi-view editing. In addition, we design a selective editing pipeline that allows users to choose preferred results from multiple candidates, offering greater flexibility and user control. Extensive experiments show that CoreEditor produces high-quality, 3D-consistent edits with sharper details, significantly outperforming prior methods.
title CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.11603