Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Ye, Sun, Zeyi, Wu, Tong, Wang, Jiaqi, Liu, Ziwei, Wetzstein, Gordon, Lin, Dahua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909209755385856
author Fang, Ye
Sun, Zeyi
Wu, Tong
Wang, Jiaqi
Liu, Ziwei
Wetzstein, Gordon
Lin, Dahua
author_facet Fang, Ye
Sun, Zeyi
Wu, Tong
Wang, Jiaqi
Liu, Ziwei
Wetzstein, Gordon
Lin, Dahua
contents Physically realistic materials are pivotal in augmenting the realism of 3D assets across various applications and lighting conditions. However, existing 3D assets and generative models often lack authentic material properties. Manual assignment of materials using graphic software is a tedious and time-consuming task. In this paper, we exploit advancements in Multimodal Large Language Models (MLLMs), particularly GPT-4V, to present a novel approach, Make-it-Real: 1) We demonstrate that GPT-4V can effectively recognize and describe materials, allowing the construction of a detailed material library. 2) Utilizing a combination of visual cues and hierarchical text prompts, GPT-4V precisely identifies and aligns materials with the corresponding components of 3D objects. 3) The correctly matched materials are then meticulously applied as reference for the new SVBRDF material generation according to the original albedo map, significantly enhancing their visual authenticity. Make-it-Real offers a streamlined integration into the 3D content creation workflow, showcasing its utility as an essential tool for developers of 3D assets.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16829
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
Fang, Ye
Sun, Zeyi
Wu, Tong
Wang, Jiaqi
Liu, Ziwei
Wetzstein, Gordon
Lin, Dahua
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Physically realistic materials are pivotal in augmenting the realism of 3D assets across various applications and lighting conditions. However, existing 3D assets and generative models often lack authentic material properties. Manual assignment of materials using graphic software is a tedious and time-consuming task. In this paper, we exploit advancements in Multimodal Large Language Models (MLLMs), particularly GPT-4V, to present a novel approach, Make-it-Real: 1) We demonstrate that GPT-4V can effectively recognize and describe materials, allowing the construction of a detailed material library. 2) Utilizing a combination of visual cues and hierarchical text prompts, GPT-4V precisely identifies and aligns materials with the corresponding components of 3D objects. 3) The correctly matched materials are then meticulously applied as reference for the new SVBRDF material generation according to the original albedo map, significantly enhancing their visual authenticity. Make-it-Real offers a streamlined integration into the 3D content creation workflow, showcasing its utility as an essential tool for developers of 3D assets.
title Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2404.16829