GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: von Lützow, Nicolas, Rössle, Barbara, Schmid, Katharina, Nießner, Matthias
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911548531802112
author von Lützow, Nicolas
Rössle, Barbara
Schmid, Katharina
Nießner, Matthias
author_facet von Lützow, Nicolas
Rössle, Barbara
Schmid, Katharina
Nießner, Matthias
contents Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer-based model that directly generates 3D Gaussians via next-token prediction, thus facilitating full 3D scene generation. We first compress Gaussian primitives into a discrete latent grid using a sparse 3D convolutional autoencoder with vector quantization. The resulting tokens are serialized and modeled using a causal transformer with 3D rotary positional embedding, enabling sequential generation of spatial structure and appearance. Unlike diffusion-based methods that refine scenes holistically, our formulation constructs scenes step-by-step, naturally supporting completion, outpainting, controllable sampling via temperature, and flexible generation horizons. This formulation leverages the compositional inductive biases and scalability of autoregressive modeling while operating on explicit representations compatible with modern neural rendering pipelines, positioning autoregressive transformers as a complementary paradigm for controllable and context-aware 3D generation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26661
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
von Lützow, Nicolas
Rössle, Barbara
Schmid, Katharina
Nießner, Matthias
Computer Vision and Pattern Recognition
Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer-based model that directly generates 3D Gaussians via next-token prediction, thus facilitating full 3D scene generation. We first compress Gaussian primitives into a discrete latent grid using a sparse 3D convolutional autoencoder with vector quantization. The resulting tokens are serialized and modeled using a causal transformer with 3D rotary positional embedding, enabling sequential generation of spatial structure and appearance. Unlike diffusion-based methods that refine scenes holistically, our formulation constructs scenes step-by-step, naturally supporting completion, outpainting, controllable sampling via temperature, and flexible generation horizons. This formulation leverages the compositional inductive biases and scalability of autoregressive modeling while operating on explicit representations compatible with modern neural rendering pipelines, positioning autoregressive transformers as a complementary paradigm for controllable and context-aware 3D generation.
title GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.26661