Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Zhiyuan, Liu, Jiuming, Chen, Yuxin, Tomizuka, Masayoshi, Xu, Chenfeng, Peng, Chensheng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918448619061248
author Xu, Zhiyuan
Liu, Jiuming
Chen, Yuxin
Tomizuka, Masayoshi
Xu, Chenfeng
Peng, Chensheng
author_facet Xu, Zhiyuan
Liu, Jiuming
Chen, Yuxin
Tomizuka, Masayoshi
Xu, Chenfeng
Peng, Chensheng
contents We present SparseGen, a novel framework for efficient image-to-3D generation, which exhibits low input-view bias while being significantly faster. Unlike traditional approaches that rely on dense volumetric grids, triplanes, or pixel-aligned primitives, we model scenes with a compact sparse set of learned 3D anchor queries and a learned expansion operator that decodes each transformed query into a small local set of 3D Gaussian primitives. Trained under a rectified-flow reconstruction objective without 3D supervision, our model learns to allocate representation capacity where geometry and appearance matter, achieving significant reductions in memory and inference time while preserving multi-view fidelity. We introduce quantitative measures of input-view bias and utilization to show that sparse queries reduce overfitting to conditioning views while being representationally efficient. Our results argue that sparse set-latent expansion is a principled, practical alternative for efficient 3D generative modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13905
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias
Xu, Zhiyuan
Liu, Jiuming
Chen, Yuxin
Tomizuka, Masayoshi
Xu, Chenfeng
Peng, Chensheng
Computer Vision and Pattern Recognition
We present SparseGen, a novel framework for efficient image-to-3D generation, which exhibits low input-view bias while being significantly faster. Unlike traditional approaches that rely on dense volumetric grids, triplanes, or pixel-aligned primitives, we model scenes with a compact sparse set of learned 3D anchor queries and a learned expansion operator that decodes each transformed query into a small local set of 3D Gaussian primitives. Trained under a rectified-flow reconstruction objective without 3D supervision, our model learns to allocate representation capacity where geometry and appearance matter, achieving significant reductions in memory and inference time while preserving multi-view fidelity. We introduce quantitative measures of input-view bias and utilization to show that sparse queries reduce overfitting to conditioning views while being representationally efficient. Our results argue that sparse set-latent expansion is a principled, practical alternative for efficient 3D generative modeling.
title Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.13905