OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zixian, Chen, Zhaoxi, Pan, Liang, Liu, Ziwei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915839012241408
author Liu, Zixian
Chen, Zhaoxi
Pan, Liang
Liu, Ziwei
author_facet Liu, Zixian
Chen, Zhaoxi
Pan, Liang
Liu, Ziwei
contents In recent years, researchers have increasingly been interested in how to enable Multimodal Large Language Models (MLLM) to possess spatial understanding and reasoning capabilities. However, most existing methods overlook the importance of the ability to continuously work in an ever-changing world, and lack the possibility of deployment on embodied systems in real-world environments. In this work, we introduce OnlineSI, a framework that can continuously improve its spatial understanding of its surroundings given a video stream. Our core idea is to maintain a finite spatial memory to retain past observations, ensuring the size of the spatial memory does not increase as the input accumulates. We further integrate 3D point cloud information with semantic information, helping MLLM to better locate and identify objects in the scene. To evaluate our method, we introduce the Fuzzy $F_1$-Score to mitigate ambiguity, and test our method on two representative datasets. Experiments demonstrate the effectiveness of our method, paving the way towards real-world embodied systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16538
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
Liu, Zixian
Chen, Zhaoxi
Pan, Liang
Liu, Ziwei
Computer Vision and Pattern Recognition
In recent years, researchers have increasingly been interested in how to enable Multimodal Large Language Models (MLLM) to possess spatial understanding and reasoning capabilities. However, most existing methods overlook the importance of the ability to continuously work in an ever-changing world, and lack the possibility of deployment on embodied systems in real-world environments. In this work, we introduce OnlineSI, a framework that can continuously improve its spatial understanding of its surroundings given a video stream. Our core idea is to maintain a finite spatial memory to retain past observations, ensuring the size of the spatial memory does not increase as the input accumulates. We further integrate 3D point cloud information with semantic information, helping MLLM to better locate and identify objects in the scene. To evaluate our method, we introduce the Fuzzy $F_1$-Score to mitigate ambiguity, and test our method on two representative datasets. Experiments demonstrate the effectiveness of our method, paving the way towards real-world embodied systems.
title OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.16538