M3: 3D-Spatial MultiModal Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Xueyan, Song, Yuchen, Qiu, Ri-Zhao, Peng, Xuanbin, Ye, Jianglong, Liu, Sifei, Wang, Xiaolong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GraspSplats: Efficient Manipulation with 3D Feature Splatting
by: Ji, Mazeyu, et al.
Published: (2024)
by: Ji, Mazeyu, et al.
Published: (2024)
WildLMa: Long Horizon Loco-Manipulation in the Wild
by: Qiu, Ri-Zhao, et al.
Published: (2024)
by: Qiu, Ri-Zhao, et al.
Published: (2024)
Dex1B: Learning with 1B Demonstrations for Dexterous Manipulation
by: Ye, Jianglong, et al.
Published: (2025)
by: Ye, Jianglong, et al.
Published: (2025)
Learning Generalizable Feature Fields for Mobile Manipulation
by: Qiu, Ri-Zhao, et al.
Published: (2024)
by: Qiu, Ri-Zhao, et al.
Published: (2024)
GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation
by: Jiang, Guangqi, et al.
Published: (2025)
by: Jiang, Guangqi, et al.
Published: (2025)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
Learning Continuous Grasping Function with a Dexterous Hand from Human Demonstrations
by: Ye, Jianglong, et al.
Published: (2022)
by: Ye, Jianglong, et al.
Published: (2022)
VLMFusionOcc3D: VLM Assisted Multi-Modal 3D Semantic Occupancy Prediction
by: Doruk, A. Enes, et al.
Published: (2026)
by: Doruk, A. Enes, et al.
Published: (2026)
Long-Tailed 3D Detection via Multi-Modal Fusion
by: Ma, Yechi, et al.
Published: (2023)
by: Ma, Yechi, et al.
Published: (2023)
Re-M3Dr: Rebalanced MultiModal Mean Deviation Regression
by: Yin, Haojie, et al.
Published: (2026)
by: Yin, Haojie, et al.
Published: (2026)
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
by: Lu, Yiren, et al.
Published: (2026)
by: Lu, Yiren, et al.
Published: (2026)
BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2025)
by: Zhang, Guowen, et al.
Published: (2025)
Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
by: Sastry, Srikumar, et al.
Published: (2025)
by: Sastry, Srikumar, et al.
Published: (2025)
Fusion-Poly: A Polyhedral Framework Based on Spatial-Temporal Fusion for 3D Multi-Object Tracking
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection
by: Chen, Zhili, et al.
Published: (2024)
by: Chen, Zhili, et al.
Published: (2024)
DNAct: Diffusion Guided Multi-Task 3D Policy Learning
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
OccCylindrical: Multi-Modal Fusion with Cylindrical Representation for 3D Semantic Occupancy Prediction
by: Ming, Zhenxing, et al.
Published: (2025)
by: Ming, Zhenxing, et al.
Published: (2025)
LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency
by: Yan, Weilong, et al.
Published: (2026)
by: Yan, Weilong, et al.
Published: (2026)
GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
by: Ze, Yanjie, et al.
Published: (2023)
by: Ze, Yanjie, et al.
Published: (2023)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
by: Yang, Yuncong, et al.
Published: (2024)
by: Yang, Yuncong, et al.
Published: (2024)
Visual Whole-Body Control for Legged Loco-Manipulation
by: Liu, Minghuan, et al.
Published: (2024)
by: Liu, Minghuan, et al.
Published: (2024)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
DBA-Fusion: Tightly Integrating Deep Dense Visual Bundle Adjustment with Multiple Sensors for Large-Scale Localization and Mapping
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
SF-Loc: A Visual Mapping and Geo-Localization System based on Sparse Visual Structure Frames
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
PNAS-MOT: Multi-Modal Object Tracking with Pareto Neural Architecture Search
by: Peng, Chensheng, et al.
Published: (2024)
by: Peng, Chensheng, et al.
Published: (2024)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
MultiModal Action Conditioned Video Generation
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
MultiModal Fine-tuning with Synthetic Captions
by: Enomoto, Shohei, et al.
Published: (2026)
by: Enomoto, Shohei, et al.
Published: (2026)
CS3D: An Efficient Facial Expression Recognition via Event Vision
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data
by: Cai, Xiongyi, et al.
Published: (2025)
by: Cai, Xiongyi, et al.
Published: (2025)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor Modalities
by: Wang, Shiming, et al.
Published: (2023)
by: Wang, Shiming, et al.
Published: (2023)
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
by: Kong, Chenqi, et al.
Published: (2023)
by: Kong, Chenqi, et al.
Published: (2023)
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
by: Singh, Binod, et al.
Published: (2025)
by: Singh, Binod, et al.
Published: (2025)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Memory Over Maps: 3D Object Localization Without Reconstruction
by: Zhou, Rui, et al.
Published: (2026)
by: Zhou, Rui, et al.
Published: (2026)
Similar Items
-
GraspSplats: Efficient Manipulation with 3D Feature Splatting
by: Ji, Mazeyu, et al.
Published: (2024) -
WildLMa: Long Horizon Loco-Manipulation in the Wild
by: Qiu, Ri-Zhao, et al.
Published: (2024) -
Dex1B: Learning with 1B Demonstrations for Dexterous Manipulation
by: Ye, Jianglong, et al.
Published: (2025) -
Learning Generalizable Feature Fields for Mobile Manipulation
by: Qiu, Ri-Zhao, et al.
Published: (2024) -
GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation
by: Jiang, Guangqi, et al.
Published: (2025)