Spice-E : Structural Priors in 3D Diffusion using Cross-Entity Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sella, Etai, Fiebelman, Gal, Atia, Noam, Averbuch-Elor, Hadar
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914788111548416
author Sella, Etai
Fiebelman, Gal
Atia, Noam
Averbuch-Elor, Hadar
author_facet Sella, Etai
Fiebelman, Gal
Atia, Noam
Averbuch-Elor, Hadar
contents We are witnessing rapid progress in automatically generating and manipulating 3D assets due to the availability of pretrained text-image diffusion models. However, time-consuming optimization procedures are required for synthesizing each sample, hindering their potential for democratizing 3D content creation. Conversely, 3D diffusion models now train on million-scale 3D datasets, yielding high-quality text-conditional 3D samples within seconds. In this work, we present Spice-E - a neural network that adds structural guidance to 3D diffusion models, extending their usage beyond text-conditional generation. At its core, our framework introduces a cross-entity attention mechanism that allows for multiple entities (in particular, paired input and guidance 3D shapes) to interact via their internal representations within the denoising network. We utilize this mechanism for learning task-specific structural priors in 3D diffusion models from auxiliary guidance shapes. We show that our approach supports a variety of applications, including 3D stylization, semantic shape editing and text-conditional abstraction-to-3D, which transforms primitive-based abstractions into highly-expressive shapes. Extensive experiments demonstrate that Spice-E achieves SOTA performance over these tasks while often being considerably faster than alternative methods. Importantly, this is accomplished without tailoring our approach for any specific task.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17834
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Spice-E : Structural Priors in 3D Diffusion using Cross-Entity Attention
Sella, Etai
Fiebelman, Gal
Atia, Noam
Averbuch-Elor, Hadar
Computer Vision and Pattern Recognition
Graphics
We are witnessing rapid progress in automatically generating and manipulating 3D assets due to the availability of pretrained text-image diffusion models. However, time-consuming optimization procedures are required for synthesizing each sample, hindering their potential for democratizing 3D content creation. Conversely, 3D diffusion models now train on million-scale 3D datasets, yielding high-quality text-conditional 3D samples within seconds. In this work, we present Spice-E - a neural network that adds structural guidance to 3D diffusion models, extending their usage beyond text-conditional generation. At its core, our framework introduces a cross-entity attention mechanism that allows for multiple entities (in particular, paired input and guidance 3D shapes) to interact via their internal representations within the denoising network. We utilize this mechanism for learning task-specific structural priors in 3D diffusion models from auxiliary guidance shapes. We show that our approach supports a variety of applications, including 3D stylization, semantic shape editing and text-conditional abstraction-to-3D, which transforms primitive-based abstractions into highly-expressive shapes. Extensive experiments demonstrate that Spice-E achieves SOTA performance over these tasks while often being considerably faster than alternative methods. Importantly, this is accomplished without tailoring our approach for any specific task.
title Spice-E : Structural Priors in 3D Diffusion using Cross-Entity Attention
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2311.17834