GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ocal, Melis, Xing, Xiaoyan, Li, Yue, Vien, Ngo Anh, Karaoglu, Sezer, Gevers, Theo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914178727411712
author Ocal, Melis
Xing, Xiaoyan
Li, Yue
Vien, Ngo Anh
Karaoglu, Sezer
Gevers, Theo
author_facet Ocal, Melis
Xing, Xiaoyan
Li, Yue
Vien, Ngo Anh
Karaoglu, Sezer
Gevers, Theo
contents 3D stylization is central to game development, virtual reality, and digital arts, where the demand for diverse assets calls for scalable methods that support fast, high-fidelity manipulation. Existing text-to-3D stylization methods typically distill from 2D image editors, requiring time-intensive per-asset optimization and exhibiting multi-view inconsistency due to the limitations of current text-to-image models, which makes them impractical for large-scale production. In this paper, we introduce GaussianBlender, a pioneering feed-forward framework for text-driven 3D stylization that performs edits instantly at inference. Our method learns structured, disentangled latent spaces with controlled information sharing for geometry and appearance from spatially-grouped 3D Gaussians. A latent diffusion model then applies text-conditioned edits on these learned representations. Comprehensive evaluations show that GaussianBlender not only delivers instant, high-fidelity, geometry-preserving, multi-view consistent stylization, but also surpasses methods that require per-instance test-time optimization - unlocking practical, democratized 3D stylization at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03683
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces
Ocal, Melis
Xing, Xiaoyan
Li, Yue
Vien, Ngo Anh
Karaoglu, Sezer
Gevers, Theo
Computer Vision and Pattern Recognition
3D stylization is central to game development, virtual reality, and digital arts, where the demand for diverse assets calls for scalable methods that support fast, high-fidelity manipulation. Existing text-to-3D stylization methods typically distill from 2D image editors, requiring time-intensive per-asset optimization and exhibiting multi-view inconsistency due to the limitations of current text-to-image models, which makes them impractical for large-scale production. In this paper, we introduce GaussianBlender, a pioneering feed-forward framework for text-driven 3D stylization that performs edits instantly at inference. Our method learns structured, disentangled latent spaces with controlled information sharing for geometry and appearance from spatially-grouped 3D Gaussians. A latent diffusion model then applies text-conditioned edits on these learned representations. Comprehensive evaluations show that GaussianBlender not only delivers instant, high-fidelity, geometry-preserving, multi-view consistent stylization, but also surpasses methods that require per-instance test-time optimization - unlocking practical, democratized 3D stylization at scale.
title GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.03683