3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Wenbo, Hong, Yining, Wang, Yanjun, Gao, Leison, Wei, Zibu, Yao, Xingcheng, Peng, Nanyun, Bitton, Yonatan, Szpektor, Idan, Chang, Kai-Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
von: Wu, Di, et al.
Veröffentlicht: (2026)
von: Wu, Di, et al.
Veröffentlicht: (2026)
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
von: Ramos, Vasco, et al.
Veröffentlicht: (2024)
von: Ramos, Vasco, et al.
Veröffentlicht: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
von: Gordon, Brian, et al.
Veröffentlicht: (2025)
von: Gordon, Brian, et al.
Veröffentlicht: (2025)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
von: Bei, Yuanchen, et al.
Veröffentlicht: (2026)
von: Bei, Yuanchen, et al.
Veröffentlicht: (2026)
Distinguishing Ignorance from Error in LLM Hallucinations
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
von: Slobodkin, Aviv, et al.
Veröffentlicht: (2025)
von: Slobodkin, Aviv, et al.
Veröffentlicht: (2025)
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
von: Bordalo, João, et al.
Veröffentlicht: (2024)
von: Bordalo, João, et al.
Veröffentlicht: (2024)
HyperMem: Hypergraph Memory for Long-Term Conversations
von: Yue, Juwei, et al.
Veröffentlicht: (2026)
von: Yue, Juwei, et al.
Veröffentlicht: (2026)
ES-Mem: Event Segmentation-Based Memory for Long-Term Dialogue Agents
von: Zou, Huhai, et al.
Veröffentlicht: (2026)
von: Zou, Huhai, et al.
Veröffentlicht: (2026)
TiMem: Temporal-Hierarchical Memory Consolidation for Long-Horizon Conversational Agents
von: Li, Kai, et al.
Veröffentlicht: (2026)
von: Li, Kai, et al.
Veröffentlicht: (2026)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
von: Yang, Yuncong, et al.
Veröffentlicht: (2024)
von: Yang, Yuncong, et al.
Veröffentlicht: (2024)
TeleMem: Building Long-Term and Multimodal Memory for Agentic AI
von: Chen, Chunliang, et al.
Veröffentlicht: (2025)
von: Chen, Chunliang, et al.
Veröffentlicht: (2025)
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
von: Hong, Yining, et al.
Veröffentlicht: (2025)
von: Hong, Yining, et al.
Veröffentlicht: (2025)
DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory
von: Qiu, Wentao, et al.
Veröffentlicht: (2026)
von: Qiu, Wentao, et al.
Veröffentlicht: (2026)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
von: Hu, Tianyu, et al.
Veröffentlicht: (2026)
von: Hu, Tianyu, et al.
Veröffentlicht: (2026)
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
MemReader: From Passive to Active Extraction for Long-Term Agent Memory
von: Kang, Jingyi, et al.
Veröffentlicht: (2026)
von: Kang, Jingyi, et al.
Veröffentlicht: (2026)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
von: Shen, Yiting, et al.
Veröffentlicht: (2026)
von: Shen, Yiting, et al.
Veröffentlicht: (2026)
DLLM Agent: See Farther, Run Faster
von: Zhen, Huiling, et al.
Veröffentlicht: (2026)
von: Zhen, Huiling, et al.
Veröffentlicht: (2026)
CarMem: Enhancing Long-Term Memory in LLM Voice Assistants through Category-Bounding
von: Kirmayr, Johannes, et al.
Veröffentlicht: (2025)
von: Kirmayr, Johannes, et al.
Veröffentlicht: (2025)
Chem3DLLM: 3D Multimodal Large Language Models for Chemistry
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
von: Shen, Zhiyu, et al.
Veröffentlicht: (2026)
von: Shen, Zhiyu, et al.
Veröffentlicht: (2026)
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
von: Ren, Xiyu, et al.
Veröffentlicht: (2026)
von: Ren, Xiyu, et al.
Veröffentlicht: (2026)
Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
von: Chhikara, Prateek, et al.
Veröffentlicht: (2025)
von: Chhikara, Prateek, et al.
Veröffentlicht: (2025)
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
von: Ramos, Vasco, et al.
Veröffentlicht: (2025)
von: Ramos, Vasco, et al.
Veröffentlicht: (2025)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios
von: Ding, Yihang, et al.
Veröffentlicht: (2026)
von: Ding, Yihang, et al.
Veröffentlicht: (2026)
TA-Mem: Tool-Augmented Autonomous Memory Retrieval for LLM in Long-Term Conversational QA
von: Yuan, Mengwei, et al.
Veröffentlicht: (2026)
von: Yuan, Mengwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024) -
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
von: Zhang, Yue, et al.
Veröffentlicht: (2026) -
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
von: Wu, Di, et al.
Veröffentlicht: (2026) -
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
von: Ramos, Vasco, et al.
Veröffentlicht: (2024) -
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)