M3: 3D-Spatial MultiModal Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Xueyan, Song, Yuchen, Qiu, Ri-Zhao, Peng, Xuanbin, Ye, Jianglong, Liu, Sifei, Wang, Xiaolong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909545878519808
author Zou, Xueyan
Song, Yuchen
Qiu, Ri-Zhao
Peng, Xuanbin
Ye, Jianglong
Liu, Sifei
Wang, Xiaolong
author_facet Zou, Xueyan
Song, Yuchen
Qiu, Ri-Zhao
Peng, Xuanbin
Ye, Jianglong
Liu, Sifei
Wang, Xiaolong
contents We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception. By integrating 3D Gaussian Splatting techniques with foundation models, M3 builds a multimodal memory capable of rendering feature representations across granularities, encompassing a wide range of knowledge. In our exploration, we identify two key challenges in previous works on feature splatting: (1) computational constraints in storing high-dimensional features for each Gaussian primitive, and (2) misalignment or information loss between distilled features and foundation model features. To address these challenges, we propose M3 with key components of principal scene components and Gaussian memory attention, enabling efficient training and inference. To validate M3, we conduct comprehensive quantitative evaluations of feature similarity and downstream tasks, as well as qualitative visualizations to highlight the pixel trace of Gaussian memory attention. Our approach encompasses a diverse range of foundation models, including vision-language models (VLMs), perception models, and large multimodal and language models (LMMs/LLMs). Furthermore, to demonstrate real-world applicability, we deploy M3's feature field in indoor scenes on a quadruped robot. Notably, we claim that M3 is the first work to address the core compression challenges in 3D feature distillation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M3: 3D-Spatial MultiModal Memory
Zou, Xueyan
Song, Yuchen
Qiu, Ri-Zhao
Peng, Xuanbin
Ye, Jianglong
Liu, Sifei
Wang, Xiaolong
Computer Vision and Pattern Recognition
Robotics
We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception. By integrating 3D Gaussian Splatting techniques with foundation models, M3 builds a multimodal memory capable of rendering feature representations across granularities, encompassing a wide range of knowledge. In our exploration, we identify two key challenges in previous works on feature splatting: (1) computational constraints in storing high-dimensional features for each Gaussian primitive, and (2) misalignment or information loss between distilled features and foundation model features. To address these challenges, we propose M3 with key components of principal scene components and Gaussian memory attention, enabling efficient training and inference. To validate M3, we conduct comprehensive quantitative evaluations of feature similarity and downstream tasks, as well as qualitative visualizations to highlight the pixel trace of Gaussian memory attention. Our approach encompasses a diverse range of foundation models, including vision-language models (VLMs), perception models, and large multimodal and language models (LMMs/LLMs). Furthermore, to demonstrate real-world applicability, we deploy M3's feature field in indoor scenes on a quadruped robot. Notably, we claim that M3 is the first work to address the core compression challenges in 3D feature distillation.
title M3: 3D-Spatial MultiModal Memory
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.16413