Structurally Disentangled Feature Fields Distillation for 3D Understanding and Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Levy, Yoel, Shavin, David, Lang, Itai, Benaim, Sagie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916623089139712
author Levy, Yoel
Shavin, David
Lang, Itai
Benaim, Sagie
author_facet Levy, Yoel
Shavin, David
Lang, Itai
Benaim, Sagie
contents Recent work has demonstrated the ability to leverage or distill pre-trained 2D features obtained using large pre-trained 2D models into 3D features, enabling impressive 3D editing and understanding capabilities using only 2D supervision. Although impressive, models assume that 3D features are captured using a single feature field and often make a simplifying assumption that features are view-independent. In this work, we propose instead to capture 3D features using multiple disentangled feature fields that capture different structural components of 3D features involving view-dependent and view-independent components, which can be learned from 2D feature supervision only. Subsequently, each element can be controlled in isolation, enabling semantic and structural understanding and editing capabilities. For instance, using a user click, one can segment 3D features corresponding to a given object and then segment, edit, or remove their view-dependent (reflective) properties. We evaluate our approach on the task of 3D segmentation and demonstrate a set of novel understanding and editing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structurally Disentangled Feature Fields Distillation for 3D Understanding and Editing
Levy, Yoel
Shavin, David
Lang, Itai
Benaim, Sagie
Computer Vision and Pattern Recognition
Recent work has demonstrated the ability to leverage or distill pre-trained 2D features obtained using large pre-trained 2D models into 3D features, enabling impressive 3D editing and understanding capabilities using only 2D supervision. Although impressive, models assume that 3D features are captured using a single feature field and often make a simplifying assumption that features are view-independent. In this work, we propose instead to capture 3D features using multiple disentangled feature fields that capture different structural components of 3D features involving view-dependent and view-independent components, which can be learned from 2D feature supervision only. Subsequently, each element can be controlled in isolation, enabling semantic and structural understanding and editing capabilities. For instance, using a user click, one can segment 3D features corresponding to a given object and then segment, edit, or remove their view-dependent (reflective) properties. We evaluate our approach on the task of 3D segmentation and demonstrate a set of novel understanding and editing tasks.
title Structurally Disentangled Feature Fields Distillation for 3D Understanding and Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.14789