Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Dingning, Huang, Xiaoshui, Hou, Yuenan, Wang, Zhihui, Yin, Zhenfei, Gong, Yongshun, Gao, Peng, Ouyang, Wanli
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913223452655616
author Liu, Dingning
Huang, Xiaoshui
Hou, Yuenan
Wang, Zhihui
Yin, Zhenfei
Gong, Yongshun
Gao, Peng
Ouyang, Wanli
author_facet Liu, Dingning
Huang, Xiaoshui
Hou, Yuenan
Wang, Zhihui
Yin, Zhenfei
Gong, Yongshun
Gao, Peng
Ouyang, Wanli
contents In this paper, we introduce Uni3D-LLM, a unified framework that leverages a Large Language Model (LLM) to integrate tasks of 3D perception, generation, and editing within point cloud scenes. This framework empowers users to effortlessly generate and modify objects at specified locations within a scene, guided by the versatility of natural language descriptions. Uni3D-LLM harnesses the expressive power of natural language to allow for precise command over the generation and editing of 3D objects, thereby significantly enhancing operational flexibility and controllability. By mapping point cloud into the unified representation space, Uni3D-LLM achieves cross-application functionality, enabling the seamless execution of a wide array of tasks, ranging from the accurate instantiation of 3D objects to the diverse requirements of interactive design. Through a comprehensive suite of rigorous experiments, the efficacy of Uni3D-LLM in the comprehension, generation, and editing of point cloud has been validated. Additionally, we have assessed the impact of integrating a point cloud perception module on the generation and editing processes, confirming the substantial potential of our approach for practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03327
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
Liu, Dingning
Huang, Xiaoshui
Hou, Yuenan
Wang, Zhihui
Yin, Zhenfei
Gong, Yongshun
Gao, Peng
Ouyang, Wanli
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
In this paper, we introduce Uni3D-LLM, a unified framework that leverages a Large Language Model (LLM) to integrate tasks of 3D perception, generation, and editing within point cloud scenes. This framework empowers users to effortlessly generate and modify objects at specified locations within a scene, guided by the versatility of natural language descriptions. Uni3D-LLM harnesses the expressive power of natural language to allow for precise command over the generation and editing of 3D objects, thereby significantly enhancing operational flexibility and controllability. By mapping point cloud into the unified representation space, Uni3D-LLM achieves cross-application functionality, enabling the seamless execution of a wide array of tasks, ranging from the accurate instantiation of 3D objects to the diverse requirements of interactive design. Through a comprehensive suite of rigorous experiments, the efficacy of Uni3D-LLM in the comprehension, generation, and editing of point cloud has been validated. Additionally, we have assessed the impact of integrating a point cloud perception module on the generation and editing processes, confirming the substantial potential of our approach for practical applications.
title Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2402.03327