ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Zhenglin, Ma, Fan, Xia, Xiaobo, Fan, Hehe, Yang, Yi, Chua, Tat-Seng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908679315390464
author Zhou, Zhenglin
Ma, Fan
Xia, Xiaobo
Fan, Hehe
Yang, Yi
Chua, Tat-Seng
author_facet Zhou, Zhenglin
Ma, Fan
Xia, Xiaobo
Fan, Hehe
Yang, Yi
Chua, Tat-Seng
contents We explore inference-time scaling in text-guided 3D diffusion models to enhance generative quality without additional training. To this end, we introduce ITS3D, a framework that formulates the task as an optimization problem to identify the most effective Gaussian noise input. The framework is driven by a verifier-guided search algorithm, where the search algorithm iteratively refines noise candidates based on verifier feedback. To address the inherent challenges of 3D generation, we introduce three techniques for improved stability, efficiency, and exploration capability. 1) Gaussian normalization is applied to stabilize the search process. It corrects distribution shifts when noise candidates deviate from a standard Gaussian distribution during iterative updates. 2) The high-dimensional nature of the 3D search space increases computational complexity. To mitigate this, a singular value decomposition-based compression technique is employed to reduce dimensionality while preserving effective search directions. 3) To further prevent convergence to suboptimal local minima, a singular space reset mechanism dynamically updates the search space based on diversity measures. Extensive experiments demonstrate that ITS3D enhances text-to-3D generation quality, which shows the potential of computationally efficient search methods in generative processes. The source code is available at https://github.com/ZhenglinZhou/ITS3D.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
Zhou, Zhenglin
Ma, Fan
Xia, Xiaobo
Fan, Hehe
Yang, Yi
Chua, Tat-Seng
Computer Vision and Pattern Recognition
We explore inference-time scaling in text-guided 3D diffusion models to enhance generative quality without additional training. To this end, we introduce ITS3D, a framework that formulates the task as an optimization problem to identify the most effective Gaussian noise input. The framework is driven by a verifier-guided search algorithm, where the search algorithm iteratively refines noise candidates based on verifier feedback. To address the inherent challenges of 3D generation, we introduce three techniques for improved stability, efficiency, and exploration capability. 1) Gaussian normalization is applied to stabilize the search process. It corrects distribution shifts when noise candidates deviate from a standard Gaussian distribution during iterative updates. 2) The high-dimensional nature of the 3D search space increases computational complexity. To mitigate this, a singular value decomposition-based compression technique is employed to reduce dimensionality while preserving effective search directions. 3) To further prevent convergence to suboptimal local minima, a singular space reset mechanism dynamically updates the search space based on diversity measures. Extensive experiments demonstrate that ITS3D enhances text-to-3D generation quality, which shows the potential of computationally efficient search methods in generative processes. The source code is available at https://github.com/ZhenglinZhou/ITS3D.
title ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22456