Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dutt, Niladri Shekhar, Muralikrishnan, Sanjeev, Mitra, Niloy J.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910396947890176
author Dutt, Niladri Shekhar
Muralikrishnan, Sanjeev
Mitra, Niloy J.
author_facet Dutt, Niladri Shekhar
Muralikrishnan, Sanjeev
Mitra, Niloy J.
contents We present Diff3F as a simple, robust, and class-agnostic feature descriptor that can be computed for untextured input shapes (meshes or point clouds). Our method distills diffusion features from image foundational models onto input shapes. Specifically, we use the input shapes to produce depth and normal maps as guidance for conditional image synthesis. In the process, we produce (diffusion) features in 2D that we subsequently lift and aggregate on the original surface. Our key observation is that even if the conditional image generations obtained from multi-view rendering of the input shapes are inconsistent, the associated image features are robust and, hence, can be directly aggregated across views. This produces semantic features on the input shapes, without requiring additional data or training. We perform extensive experiments on multiple benchmarks (SHREC'19, SHREC'20, FAUST, and TOSCA) and demonstrate that our features, being semantic instead of geometric, produce reliable correspondence across both isometric and non-isometrically related shape families. Code is available via the project page at https://diff3f.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2311_17024
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features
Dutt, Niladri Shekhar
Muralikrishnan, Sanjeev
Mitra, Niloy J.
Computer Vision and Pattern Recognition
Graphics
We present Diff3F as a simple, robust, and class-agnostic feature descriptor that can be computed for untextured input shapes (meshes or point clouds). Our method distills diffusion features from image foundational models onto input shapes. Specifically, we use the input shapes to produce depth and normal maps as guidance for conditional image synthesis. In the process, we produce (diffusion) features in 2D that we subsequently lift and aggregate on the original surface. Our key observation is that even if the conditional image generations obtained from multi-view rendering of the input shapes are inconsistent, the associated image features are robust and, hence, can be directly aggregated across views. This produces semantic features on the input shapes, without requiring additional data or training. We perform extensive experiments on multiple benchmarks (SHREC'19, SHREC'20, FAUST, and TOSCA) and demonstrate that our features, being semantic instead of geometric, produce reliable correspondence across both isometric and non-isometrically related shape families. Code is available via the project page at https://diff3f.github.io/
title Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2311.17024