MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Jiaxin, Chen, Runnan, Li, Ziwen, Gao, Zhengqing, He, Xiao, Guo, Yandong, Gong, Mingming, Liu, Tongliang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911254899064832
author Huang, Jiaxin
Chen, Runnan
Li, Ziwen
Gao, Zhengqing
He, Xiao
Guo, Yandong
Gong, Mingming
Liu, Tongliang
author_facet Huang, Jiaxin
Chen, Runnan
Li, Ziwen
Gao, Zhengqing
He, Xiao
Guo, Yandong
Gong, Mingming
Liu, Tongliang
contents Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning segmentation, adapting these capabilities to 3D scenes remains underexplored. In this paper, we introduce MLLM-For3D, a simple yet effective framework that transfers knowledge from 2D MLLMs to 3D scene understanding. Specifically, we utilize MLLMs to generate multi-view pseudo segmentation masks and corresponding text embeddings, then unproject 2D masks into 3D space and align them with the text embeddings. The primary challenge lies in the absence of 3D context and spatial consistency across multiple views, causing the model to hallucinate objects that do not exist and fail to target objects consistently. Training the 3D model with such irrelevant objects leads to performance degradation. To address this, we introduce a spatial consistency strategy to enforce that segmentation masks remain coherent in the 3D space, effectively capturing the geometry of the scene. Moreover, we develop a Token-for-Query approach for multimodal semantic alignment, enabling consistent identification of the same object across different views. Extensive evaluations on various challenging indoor scene benchmarks demonstrate that, even without any labeled 3D training data, MLLM-For3D outperforms existing 3D reasoning segmentation methods, effectively interpreting user intent, understanding 3D scenes, and reasoning about spatial relationships.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation
Huang, Jiaxin
Chen, Runnan
Li, Ziwen
Gao, Zhengqing
He, Xiao
Guo, Yandong
Gong, Mingming
Liu, Tongliang
Computer Vision and Pattern Recognition
Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning segmentation, adapting these capabilities to 3D scenes remains underexplored. In this paper, we introduce MLLM-For3D, a simple yet effective framework that transfers knowledge from 2D MLLMs to 3D scene understanding. Specifically, we utilize MLLMs to generate multi-view pseudo segmentation masks and corresponding text embeddings, then unproject 2D masks into 3D space and align them with the text embeddings. The primary challenge lies in the absence of 3D context and spatial consistency across multiple views, causing the model to hallucinate objects that do not exist and fail to target objects consistently. Training the 3D model with such irrelevant objects leads to performance degradation. To address this, we introduce a spatial consistency strategy to enforce that segmentation masks remain coherent in the 3D space, effectively capturing the geometry of the scene. Moreover, we develop a Token-for-Query approach for multimodal semantic alignment, enabling consistent identification of the same object across different views. Extensive evaluations on various challenging indoor scene benchmarks demonstrate that, even without any labeled 3D training data, MLLM-For3D outperforms existing 3D reasoning segmentation methods, effectively interpreting user intent, understanding 3D scenes, and reasoning about spatial relationships.
title MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18135