Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shuang, Lin, Youtian, Zhang, Feihu, Zeng, Yifei, Yang, Yikang, Bao, Yajie, Qian, Jiachen, Zhu, Siyu, Cao, Xun, Torr, Philip, Yao, Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915304920055808
author Wu, Shuang
Lin, Youtian
Zhang, Feihu
Zeng, Yifei
Yang, Yikang
Bao, Yajie
Qian, Jiachen
Zhu, Siyu
Cao, Xun
Torr, Philip
Yao, Yao
author_facet Wu, Shuang
Lin, Youtian
Zhang, Feihu
Zeng, Yifei
Yang, Yikang
Bao, Yajie
Qian, Jiachen
Zhu, Siyu
Cao, Xun
Torr, Philip
Yao, Yao
contents Generating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dramatically reduced training costs. Our key innovation is the Spatial Sparse Attention (SSA) mechanism, which greatly enhances the efficiency of Diffusion Transformer (DiT) computations on sparse volumetric data. SSA allows the model to effectively process large token sets within sparse volumes, substantially reducing computational overhead and achieving a 3.9x speedup in the forward pass and a 9.6x speedup in the backward pass. Our framework also includes a variational autoencoder (VAE) that maintains a consistent sparse volumetric format across input, latent, and output stages. Compared to previous methods with heterogeneous representations in 3D VAE, this unified design significantly improves training efficiency and stability. Our model is trained on public available datasets, and experiments demonstrate that Direct3D-S2 not only surpasses state-of-the-art methods in generation quality and efficiency, but also enables training at 1024 resolution using only 8 GPUs, a task typically requiring at least 32 GPUs for volumetric representations at 256 resolution, thus making gigascale 3D generation both practical and accessible. Project page: https://www.neural4d.com/research/direct3d-s2.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17412
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
Wu, Shuang
Lin, Youtian
Zhang, Feihu
Zeng, Yifei
Yang, Yikang
Bao, Yajie
Qian, Jiachen
Zhu, Siyu
Cao, Xun
Torr, Philip
Yao, Yao
Computer Vision and Pattern Recognition
Generating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dramatically reduced training costs. Our key innovation is the Spatial Sparse Attention (SSA) mechanism, which greatly enhances the efficiency of Diffusion Transformer (DiT) computations on sparse volumetric data. SSA allows the model to effectively process large token sets within sparse volumes, substantially reducing computational overhead and achieving a 3.9x speedup in the forward pass and a 9.6x speedup in the backward pass. Our framework also includes a variational autoencoder (VAE) that maintains a consistent sparse volumetric format across input, latent, and output stages. Compared to previous methods with heterogeneous representations in 3D VAE, this unified design significantly improves training efficiency and stability. Our model is trained on public available datasets, and experiments demonstrate that Direct3D-S2 not only surpasses state-of-the-art methods in generation quality and efficiency, but also enables training at 1024 resolution using only 8 GPUs, a task typically requiring at least 32 GPUs for volumetric representations at 256 resolution, thus making gigascale 3D generation both practical and accessible. Project page: https://www.neural4d.com/research/direct3d-s2.
title Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17412