ConTEXTure: Consistent Multiview Images to Texture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahn, Jaehoon, Cho, Sumin, Jung, Harim, Hong, Kibeom, Ban, Seonghoon, Jung, Moon-Ryul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929421197246464
author Ahn, Jaehoon
Cho, Sumin
Jung, Harim
Hong, Kibeom
Ban, Seonghoon
Jung, Moon-Ryul
author_facet Ahn, Jaehoon
Cho, Sumin
Jung, Harim
Hong, Kibeom
Ban, Seonghoon
Jung, Moon-Ryul
contents We introduce ConTEXTure, a generative network designed to create a texture map/atlas for a given 3D mesh using images from multiple viewpoints. The process begins with generating a front-view image from a text prompt, such as 'Napoleon, front view', describing the 3D mesh. Additional images from different viewpoints are derived from this front-view image and camera poses relative to it. ConTEXTure builds upon the TEXTure network, which uses text prompts for six viewpoints (e.g., 'Napoleon, front view', 'Napoleon, left view', etc.). However, TEXTure often generates images for non-front viewpoints that do not accurately represent those viewpoints.To address this issue, we employ Zero123++, which generates multiple view-consistent images for the six specified viewpoints simultaneously, conditioned on the initial front-view image and the depth maps of the mesh for the six viewpoints. By utilizing these view-consistent images, ConTEXTure learns the texture atlas from all viewpoint images concurrently, unlike previous methods that do so sequentially. This approach ensures that the rendered images from various viewpoints, including back, side, bottom, and top, are free from viewpoint irregularities.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10558
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ConTEXTure: Consistent Multiview Images to Texture
Ahn, Jaehoon
Cho, Sumin
Jung, Harim
Hong, Kibeom
Ban, Seonghoon
Jung, Moon-Ryul
Computer Vision and Pattern Recognition
Machine Learning
We introduce ConTEXTure, a generative network designed to create a texture map/atlas for a given 3D mesh using images from multiple viewpoints. The process begins with generating a front-view image from a text prompt, such as 'Napoleon, front view', describing the 3D mesh. Additional images from different viewpoints are derived from this front-view image and camera poses relative to it. ConTEXTure builds upon the TEXTure network, which uses text prompts for six viewpoints (e.g., 'Napoleon, front view', 'Napoleon, left view', etc.). However, TEXTure often generates images for non-front viewpoints that do not accurately represent those viewpoints.To address this issue, we employ Zero123++, which generates multiple view-consistent images for the six specified viewpoints simultaneously, conditioned on the initial front-view image and the depth maps of the mesh for the six viewpoints. By utilizing these view-consistent images, ConTEXTure learns the texture atlas from all viewpoint images concurrently, unlike previous methods that do so sequentially. This approach ensures that the rendered images from various viewpoints, including back, side, bottom, and top, are free from viewpoint irregularities.
title ConTEXTure: Consistent Multiview Images to Texture
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2407.10558