Empowering Large Language Models with 3D Situation Awareness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Zhihao, Peng, Yibo, Ren, Jinke, Liao, Yinghong, Han, Yatong, Feng, Chun-Mei, Zhao, Hengshuang, Li, Guanbin, Cui, Shuguang, Li, Zhen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908289824980992
author Yuan, Zhihao
Peng, Yibo
Ren, Jinke
Liao, Yinghong
Han, Yatong
Feng, Chun-Mei
Zhao, Hengshuang
Li, Guanbin
Cui, Shuguang
Li, Zhen
author_facet Yuan, Zhihao
Peng, Yibo
Ren, Jinke
Liao, Yinghong
Han, Yatong
Feng, Chun-Mei
Zhao, Hengshuang
Li, Guanbin
Cui, Shuguang
Li, Zhen
contents Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions (e.g., ''left" or ''right"). However, current LLM-based methods overlook the egocentric perspective and simply use datasets from a global viewpoint. To address this issue, we propose a novel approach to automatically generate a situation-aware dataset by leveraging the scanning trajectory during data collection and utilizing Vision-Language Models (VLMs) to produce high-quality captions and question-answer pairs. Furthermore, we introduce a situation grounding module to explicitly predict the position and orientation of observer's viewpoint, thereby enabling LLMs to ground situation description in 3D scenes. We evaluate our approach on several benchmarks, demonstrating that our method effectively enhances the 3D situational awareness of LLMs while significantly expanding existing datasets and reducing manual effort.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empowering Large Language Models with 3D Situation Awareness
Yuan, Zhihao
Peng, Yibo
Ren, Jinke
Liao, Yinghong
Han, Yatong
Feng, Chun-Mei
Zhao, Hengshuang
Li, Guanbin
Cui, Shuguang
Li, Zhen
Computer Vision and Pattern Recognition
Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions (e.g., ''left" or ''right"). However, current LLM-based methods overlook the egocentric perspective and simply use datasets from a global viewpoint. To address this issue, we propose a novel approach to automatically generate a situation-aware dataset by leveraging the scanning trajectory during data collection and utilizing Vision-Language Models (VLMs) to produce high-quality captions and question-answer pairs. Furthermore, we introduce a situation grounding module to explicitly predict the position and orientation of observer's viewpoint, thereby enabling LLMs to ground situation description in 3D scenes. We evaluate our approach on several benchmarks, demonstrating that our method effectively enhances the 3D situational awareness of LLMs while significantly expanding existing datasets and reducing manual effort.
title Empowering Large Language Models with 3D Situation Awareness
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.23024