3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Jiajun, He, Tianyu, Jiang, Li, Wang, Tianyu, Dayoub, Feras, Reid, Ian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916704652623872
author Deng, Jiajun
He, Tianyu
Jiang, Li
Wang, Tianyu
Dayoub, Feras
Reid, Ian
author_facet Deng, Jiajun
He, Tianyu
Jiang, Li
Wang, Tianyu
Dayoub, Feras
Reid, Ian
contents Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we introduce 3D-LLaVA, a simple yet highly powerful 3D LMM designed to act as an intelligent assistant in comprehending, reasoning, and interacting with the 3D world. Unlike existing top-performing methods that rely on complicated pipelines-such as offline multi-view feature extraction or additional task-specific heads-3D-LLaVA adopts a minimalist design with integrated architecture and only takes point clouds as input. At the core of 3D-LLaVA is a new Omni Superpoint Transformer (OST), which integrates three functionalities: (1) a visual feature selector that converts and selects visual tokens, (2) a visual prompt encoder that embeds interactive visual prompts into the visual token space, and (3) a referring mask decoder that produces 3D masks based on text description. This versatile OST is empowered by the hybrid pretraining to obtain perception priors and leveraged as the visual connector that bridges the 3D data to the LLM. After performing unified instruction tuning, our 3D-LLaVA reports impressive results on various benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01163
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
Deng, Jiajun
He, Tianyu
Jiang, Li
Wang, Tianyu
Dayoub, Feras
Reid, Ian
Computer Vision and Pattern Recognition
Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we introduce 3D-LLaVA, a simple yet highly powerful 3D LMM designed to act as an intelligent assistant in comprehending, reasoning, and interacting with the 3D world. Unlike existing top-performing methods that rely on complicated pipelines-such as offline multi-view feature extraction or additional task-specific heads-3D-LLaVA adopts a minimalist design with integrated architecture and only takes point clouds as input. At the core of 3D-LLaVA is a new Omni Superpoint Transformer (OST), which integrates three functionalities: (1) a visual feature selector that converts and selects visual tokens, (2) a visual prompt encoder that embeds interactive visual prompts into the visual token space, and (3) a referring mask decoder that produces 3D masks based on text description. This versatile OST is empowered by the hybrid pretraining to obtain perception priors and leveraged as the visual connector that bridges the 3D data to the LLM. After performing unified instruction tuning, our 3D-LLaVA reports impressive results on various benchmarks.
title 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.01163