Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Song, Kong, Lingdong, Liu, Xiaolu, Shi, Hao, Li, Wentong, Zhu, Jianke, Hoi, Steven C. H. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not All Voxels Are Equal: Hardness-Aware Semantic Scene Completion with Self-Distillation
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving
by: Liu, Xiaolu, et al.
Published: (2026)
by: Liu, Xiaolu, et al.
Published: (2026)
Label-efficient Semantic Scene Completion with Scribble Annotations
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
ReliOcc: Towards Reliable Semantic Occupancy Prediction via Uncertainty Learning
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Traj-LIO: A Resilient Multi-LiDAR Multi-IMU State Estimator Through Sparse Gaussian Process
by: Zheng, Xin, et al.
Published: (2024)
by: Zheng, Xin, et al.
Published: (2024)
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Offboard Occupancy Refinement with Hybrid Propagation for Autonomous Driving
by: Shi, Hao, et al.
Published: (2024)
by: Shi, Hao, et al.
Published: (2024)
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
by: Xie, Shaoyuan, et al.
Published: (2025)
by: Xie, Shaoyuan, et al.
Published: (2025)
Benchmarking and Improving Bird's Eye View Perception Robustness in Autonomous Driving
by: Xie, Shaoyuan, et al.
Published: (2024)
by: Xie, Shaoyuan, et al.
Published: (2024)
Monocular Semantic Scene Completion via Masked Recurrent Networks
by: Wang, Xuzhi, et al.
Published: (2025)
by: Wang, Xuzhi, et al.
Published: (2025)
LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2025)
by: Kong, Lingdong, et al.
Published: (2025)
123D: Unifying Multi-Modal Autonomous Driving Data at Scale
by: Dauner, Daniel, et al.
Published: (2026)
by: Dauner, Daniel, et al.
Published: (2026)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
FRNet: Frustum-Range Networks for Scalable LiDAR Segmentation
by: Xu, Xiang, et al.
Published: (2023)
by: Xu, Xiang, et al.
Published: (2023)
M3: 3D-Spatial MultiModal Memory
by: Zou, Xueyan, et al.
Published: (2025)
by: Zou, Xueyan, et al.
Published: (2025)
Uncertainty-Instructed Structure Injection for Generalizable HD Map Construction
by: Liu, Xiaolu, et al.
Published: (2025)
by: Liu, Xiaolu, et al.
Published: (2025)
MGMap: Mask-Guided Learning for Online Vectorized HD Map Construction
by: Liu, Xiaolu, et al.
Published: (2024)
by: Liu, Xiaolu, et al.
Published: (2024)
HVOFusion: Incremental Mesh Reconstruction Using Hybrid Voxel Octree
by: Liu, Shaofan, et al.
Published: (2024)
by: Liu, Shaofan, et al.
Published: (2024)
3D and 4D World Modeling: A Survey
by: Kong, Lingdong, et al.
Published: (2025)
by: Kong, Lingdong, et al.
Published: (2025)
Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations
by: Xu, Xiang, et al.
Published: (2025)
by: Xu, Xiang, et al.
Published: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
Is Your LiDAR Placement Optimized for 3D Scene Understanding?
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
Graph-Based Multi-Modal Sensor Fusion for Autonomous Driving
by: Sani, Depanshu, et al.
Published: (2024)
by: Sani, Depanshu, et al.
Published: (2024)
La La LiDAR: Large-Scale Layout Generation from LiDAR Data
by: Liu, Youquan, et al.
Published: (2025)
by: Liu, Youquan, et al.
Published: (2025)
RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
by: Feng, Sicheng, et al.
Published: (2025)
by: Feng, Sicheng, et al.
Published: (2025)
An Empirical Study of Training State-of-the-Art LiDAR Segmentation Models
by: Sun, Jiahao, et al.
Published: (2024)
by: Sun, Jiahao, et al.
Published: (2024)
MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes
by: Bian, Hengwei, et al.
Published: (2024)
by: Bian, Hengwei, et al.
Published: (2024)
Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
by: Rowe, Luke, et al.
Published: (2025)
by: Rowe, Luke, et al.
Published: (2025)
M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving
by: Xu, Dongyang, et al.
Published: (2024)
by: Xu, Dongyang, et al.
Published: (2024)
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
by: Khurana, Mehar, et al.
Published: (2024)
by: Khurana, Mehar, et al.
Published: (2024)
FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational Frequencies
by: Lu, Dongyue, et al.
Published: (2024)
by: Lu, Dongyue, et al.
Published: (2024)
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025)
by: Yu, Hanxun, et al.
Published: (2025)
Enhanced Spatiotemporal Consistency for Image-to-LiDAR Data Pretraining
by: Xu, Xiang, et al.
Published: (2025)
by: Xu, Xiang, et al.
Published: (2025)
SAM4D: Segment Anything in Camera and LiDAR Streams
by: Xu, Jianyun, et al.
Published: (2025)
by: Xu, Jianyun, et al.
Published: (2025)
Multi-Space Alignments Towards Universal LiDAR Segmentation
by: Liu, Youquan, et al.
Published: (2024)
by: Liu, Youquan, et al.
Published: (2024)
Similar Items
-
Not All Voxels Are Equal: Hardness-Aware Semantic Scene Completion with Self-Distillation
by: Wang, Song, et al.
Published: (2024) -
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
by: Shi, Zhan, et al.
Published: (2025) -
DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving
by: Liu, Xiaolu, et al.
Published: (2026) -
Label-efficient Semantic Scene Completion with Scribble Annotations
by: Wang, Song, et al.
Published: (2024) -
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)