Vinedresser3D: Agentic Text-guided 3D Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chi, Yankuan, Li, Xiang, Huang, Zixuan, Rehg, James M.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908848519905280
author Chi, Yankuan
Li, Xiang
Huang, Zixuan
Rehg, James M.
author_facet Chi, Yankuan
Li, Xiang
Huang, Zixuan
Rehg, James M.
contents Text-guided 3D editing aims to modify existing 3D assets using natural-language instructions. Current methods struggle to jointly understand complex prompts, automatically localize edits in 3D, and preserve unedited content. We introduce Vinedresser3D, an agentic framework for high-quality text-guided 3D editing that operates directly in the latent space of a native 3D generative model. Given a 3D asset and an editing prompt, Vinedresser3D uses a multimodal large language model to infer rich descriptions of the original asset, identify the edit region and edit type (addition, modification, deletion), and generate decomposed structural and appearance-level text guidance. The agent then selects an informative view and applies an image editing model to obtain visual guidance. Finally, an inversion-based rectified-flow inpainting pipeline with an interleaved sampling module performs editing in the 3D latent space, enforcing prompt alignment while maintaining 3D coherence and unedited regions. Experiments on diverse 3D edits demonstrate that Vinedresser3D outperforms prior baselines in both automatic metrics and human preference studies, while enabling precise, coherent, and mask-free 3D editing.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19542
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vinedresser3D: Agentic Text-guided 3D Editing
Chi, Yankuan
Li, Xiang
Huang, Zixuan
Rehg, James M.
Computer Vision and Pattern Recognition
Text-guided 3D editing aims to modify existing 3D assets using natural-language instructions. Current methods struggle to jointly understand complex prompts, automatically localize edits in 3D, and preserve unedited content. We introduce Vinedresser3D, an agentic framework for high-quality text-guided 3D editing that operates directly in the latent space of a native 3D generative model. Given a 3D asset and an editing prompt, Vinedresser3D uses a multimodal large language model to infer rich descriptions of the original asset, identify the edit region and edit type (addition, modification, deletion), and generate decomposed structural and appearance-level text guidance. The agent then selects an informative view and applies an image editing model to obtain visual guidance. Finally, an inversion-based rectified-flow inpainting pipeline with an interleaved sampling module performs editing in the 3D latent space, enforcing prompt alignment while maintaining 3D coherence and unedited regions. Experiments on diverse 3D edits demonstrate that Vinedresser3D outperforms prior baselines in both automatic metrics and human preference studies, while enabling precise, coherent, and mask-free 3D editing.
title Vinedresser3D: Agentic Text-guided 3D Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.19542