A Scalable Attention-Based Approach for Image-to-3D Texture Mapping

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Rampini, Arianna, Madan, Kanika, Roy, Bruno, Zamani, AmirHossein, Cheung, Derek
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918136365711360
author Rampini, Arianna
Madan, Kanika
Roy, Bruno
Zamani, AmirHossein
Cheung, Derek
author_facet Rampini, Arianna
Madan, Kanika
Roy, Bruno
Zamani, AmirHossein
Cheung, Derek
contents High-quality textures are critical for realistic 3D content creation, yet existing generative methods are slow, rely on UV maps, and often fail to remain faithful to a reference image. To address these challenges, we propose a transformer-based framework that predicts a 3D texture field directly from a single image and a mesh, eliminating the need for UV mapping and differentiable rendering, and enabling faster texture generation. Our method integrates a triplane representation with depth-based backprojection losses, enabling efficient training and faster inference. Once trained, it generates high-fidelity textures in a single forward pass, requiring only 0.2s per shape. Extensive qualitative, quantitative, and user preference evaluations demonstrate that our method outperforms state-of-the-art baselines on single-image texture reconstruction in terms of both fidelity to the input image and perceptual quality, highlighting its practicality for scalable, high-quality, and controllable 3D content creation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Scalable Attention-Based Approach for Image-to-3D Texture Mapping
Rampini, Arianna
Madan, Kanika
Roy, Bruno
Zamani, AmirHossein
Cheung, Derek
Computer Vision and Pattern Recognition
Machine Learning
High-quality textures are critical for realistic 3D content creation, yet existing generative methods are slow, rely on UV maps, and often fail to remain faithful to a reference image. To address these challenges, we propose a transformer-based framework that predicts a 3D texture field directly from a single image and a mesh, eliminating the need for UV mapping and differentiable rendering, and enabling faster texture generation. Our method integrates a triplane representation with depth-based backprojection losses, enabling efficient training and faster inference. Once trained, it generates high-fidelity textures in a single forward pass, requiring only 0.2s per shape. Extensive qualitative, quantitative, and user preference evaluations demonstrate that our method outperforms state-of-the-art baselines on single-image texture reconstruction in terms of both fidelity to the input image and perceptual quality, highlighting its practicality for scalable, high-quality, and controllable 3D content creation.
title A Scalable Attention-Based Approach for Image-to-3D Texture Mapping
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.05131