3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Wenbo, Hong, Yining, Wang, Yanjun, Gao, Leison, Wei, Zibu, Yao, Xingcheng, Peng, Nanyun, Bitton, Yonatan, Szpektor, Idan, Chang, Kai-Wei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908716171788288
author Hu, Wenbo
Hong, Yining
Wang, Yanjun
Gao, Leison
Wei, Zibu
Yao, Xingcheng
Peng, Nanyun
Bitton, Yonatan
Szpektor, Idan
Chang, Kai-Wei
author_facet Hu, Wenbo
Hong, Yining
Wang, Yanjun
Gao, Leison
Wei, Zibu
Yao, Xingcheng
Peng, Nanyun
Bitton, Yonatan
Szpektor, Idan
Chang, Kai-Wei
contents Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments. We posit that part of this limitation is due to the lack of proper 3D spatial-temporal memory modeling in LLMs. To address this, we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26,000 trajectories and 2,892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments. Second, we propose 3DLLM-Mem, a novel dynamic memory management and fusion model for embodied spatial-temporal reasoning and actions in LLMs. Our model uses working memory tokens, which represents current observations, as queries to selectively attend to and fuse the most useful spatial and temporal features from episodic memory, which stores past observations and interactions. Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments. Experimental results demonstrate that 3DLLM-Mem achieves state-of-the-art performance across various tasks, outperforming the strongest baselines by 16.5% in success rate on 3DMem-Bench's most challenging in-the-wild embodied tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22657
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
Hu, Wenbo
Hong, Yining
Wang, Yanjun
Gao, Leison
Wei, Zibu
Yao, Xingcheng
Peng, Nanyun
Bitton, Yonatan
Szpektor, Idan
Chang, Kai-Wei
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments. We posit that part of this limitation is due to the lack of proper 3D spatial-temporal memory modeling in LLMs. To address this, we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26,000 trajectories and 2,892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments. Second, we propose 3DLLM-Mem, a novel dynamic memory management and fusion model for embodied spatial-temporal reasoning and actions in LLMs. Our model uses working memory tokens, which represents current observations, as queries to selectively attend to and fuse the most useful spatial and temporal features from episodic memory, which stores past observations and interactions. Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments. Experimental results demonstrate that 3DLLM-Mem achieves state-of-the-art performance across various tasks, outperforming the strongest baselines by 16.5% in success rate on 3DMem-Bench's most challenging in-the-wild embodied tasks.
title 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.22657