Instructive3D: Editing Large Reconstruction Models with Text Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kathare, Kunal, Dhiman, Ankit, Gowda, K Vikas, Aravindan, Siddharth, Monga, Shubham, Vandrotti, Basavaraja Shanthappa, Boregowda, Lokesh R
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915094599827456
author Kathare, Kunal
Dhiman, Ankit
Gowda, K Vikas
Aravindan, Siddharth
Monga, Shubham
Vandrotti, Basavaraja Shanthappa
Boregowda, Lokesh R
author_facet Kathare, Kunal
Dhiman, Ankit
Gowda, K Vikas
Aravindan, Siddharth
Monga, Shubham
Vandrotti, Basavaraja Shanthappa
Boregowda, Lokesh R
contents Transformer based methods have enabled users to create, modify, and comprehend text and image data. Recently proposed Large Reconstruction Models (LRMs) further extend this by providing the ability to generate high-quality 3D models with the help of a single object image. These models, however, lack the ability to manipulate or edit the finer details, such as adding standard design patterns or changing the color and reflectance of the generated objects, thus lacking fine-grained control that may be very helpful in domains such as augmented reality, animation and gaming. Naively training LRMs for this purpose would require generating precisely edited images and 3D object pairs, which is computationally expensive. In this paper, we propose Instructive3D, a novel LRM based model that integrates generation and fine-grained editing, through user text prompts, of 3D objects into a single model. We accomplish this by adding an adapter that performs a diffusion process conditioned on a text prompt specifying edits in the triplane latent space representation of 3D object models. Our method does not require the generation of edited 3D objects. Additionally, Instructive3D allows us to perform geometrically consistent modifications, as the edits done through user-defined text prompts are applied to the triplane latent representation thus enhancing the versatility and precision of 3D objects generated. We compare the objects generated by Instructive3D and a baseline that first generates the 3D object meshes using a standard LRM model and then edits these 3D objects using text prompts when images are provided from the Objaverse LVIS dataset. We find that Instructive3D produces qualitatively superior 3D objects with the properties specified by the edit prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instructive3D: Editing Large Reconstruction Models with Text Instructions
Kathare, Kunal
Dhiman, Ankit
Gowda, K Vikas
Aravindan, Siddharth
Monga, Shubham
Vandrotti, Basavaraja Shanthappa
Boregowda, Lokesh R
Computer Vision and Pattern Recognition
Transformer based methods have enabled users to create, modify, and comprehend text and image data. Recently proposed Large Reconstruction Models (LRMs) further extend this by providing the ability to generate high-quality 3D models with the help of a single object image. These models, however, lack the ability to manipulate or edit the finer details, such as adding standard design patterns or changing the color and reflectance of the generated objects, thus lacking fine-grained control that may be very helpful in domains such as augmented reality, animation and gaming. Naively training LRMs for this purpose would require generating precisely edited images and 3D object pairs, which is computationally expensive. In this paper, we propose Instructive3D, a novel LRM based model that integrates generation and fine-grained editing, through user text prompts, of 3D objects into a single model. We accomplish this by adding an adapter that performs a diffusion process conditioned on a text prompt specifying edits in the triplane latent space representation of 3D object models. Our method does not require the generation of edited 3D objects. Additionally, Instructive3D allows us to perform geometrically consistent modifications, as the edits done through user-defined text prompts are applied to the triplane latent representation thus enhancing the versatility and precision of 3D objects generated. We compare the objects generated by Instructive3D and a baseline that first generates the 3D object meshes using a standard LRM model and then edits these 3D objects using text prompts when images are provided from the Objaverse LVIS dataset. We find that Instructive3D produces qualitatively superior 3D objects with the properties specified by the edit prompts.
title Instructive3D: Editing Large Reconstruction Models with Text Instructions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04374