Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sella, Etai, Phung, Hao, Amiel, Nitay, Litany, Or, Patashnik, Or, Averbuch-Elor, Hadar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909000534065152
author Sella, Etai
Phung, Hao
Amiel, Nitay
Litany, Or
Patashnik, Or
Averbuch-Elor, Hadar
author_facet Sella, Etai
Phung, Hao
Amiel, Nitay
Litany, Or
Patashnik, Or
Averbuch-Elor, Hadar
contents Text-based 2D image editing models have recently reached an impressive level of maturity, motivating a growing body of work that heavily depends on these models to drive 3D edits. While effective for appearance-based modifications, such 2D-centric 3D editing pipelines often struggle with fine-grained 3D editing, where localized structural changes must be applied while strictly preserving an object's overall identity. To address this limitation, we propose Prox-E, a training-free framework that enables fine-grained 3D control through an explicit, primitive-based geometric abstraction. Our framework first abstracts an input 3D shape into a compact set of geometric primitives. A pretrained vision-language model (VLM) then edits this abstraction to specify primitive-level changes. These structural edits are subsequently used to guide a 3D generative model, enabling fine-grained, localized modifications while preserving unchanged regions of the original shape. Through extensive experiments, we demonstrate that our method consistently balances identity preservation, shape quality, and instruction fidelity more effectively than various existing approaches, including 2D-based 3D editors and training-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23774
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions
Sella, Etai
Phung, Hao
Amiel, Nitay
Litany, Or
Patashnik, Or
Averbuch-Elor, Hadar
Graphics
Text-based 2D image editing models have recently reached an impressive level of maturity, motivating a growing body of work that heavily depends on these models to drive 3D edits. While effective for appearance-based modifications, such 2D-centric 3D editing pipelines often struggle with fine-grained 3D editing, where localized structural changes must be applied while strictly preserving an object's overall identity. To address this limitation, we propose Prox-E, a training-free framework that enables fine-grained 3D control through an explicit, primitive-based geometric abstraction. Our framework first abstracts an input 3D shape into a compact set of geometric primitives. A pretrained vision-language model (VLM) then edits this abstraction to specify primitive-level changes. These structural edits are subsequently used to guide a 3D generative model, enabling fine-grained, localized modifications while preserving unchanged regions of the original shape. Through extensive experiments, we demonstrate that our method consistently balances identity preservation, shape quality, and instruction fidelity more effectively than various existing approaches, including 2D-based 3D editors and training-based methods.
title Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions
topic Graphics
url https://arxiv.org/abs/2604.23774