SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Xiangchao, Chen, Runjian, Zhang, Bo, Ye, Hancheng, Xia, Renqiu, Yuan, Jiakang, Zhou, Hongbin, Cai, Xinyu, Shi, Botian, Shao, Wenqi, Luo, Ping, Qiao, Yu, Chen, Tao, Yan, Junchi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
by: Zhang, Bo, et al.
Published: (2023)
by: Zhang, Bo, et al.
Published: (2023)
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
by: Xia, Renqiu, et al.
Published: (2023)
by: Xia, Renqiu, et al.
Published: (2023)
TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR Perception
by: Chen, Runjian, et al.
Published: (2024)
by: Chen, Runjian, et al.
Published: (2024)
CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning
by: Chen, Runjian, et al.
Published: (2024)
by: Chen, Runjian, et al.
Published: (2024)
CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving
by: Chen, Runjian, et al.
Published: (2022)
by: Chen, Runjian, et al.
Published: (2022)
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
by: Yan, Xiangchao, et al.
Published: (2025)
by: Yan, Xiangchao, et al.
Published: (2025)
SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction
by: Chen, Suzeyu, et al.
Published: (2026)
by: Chen, Suzeyu, et al.
Published: (2026)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
by: Liao, Ning, et al.
Published: (2023)
by: Liao, Ning, et al.
Published: (2023)
Temporal Overlapping Prediction: A Self-supervised Pre-training Method for LiDAR Moving Object Segmentation
by: Miao, Ziliang, et al.
Published: (2025)
by: Miao, Ziliang, et al.
Published: (2025)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data
by: Chen, Runjian, et al.
Published: (2025)
by: Chen, Runjian, et al.
Published: (2025)
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Multimodal 3D Genome Pre-training
by: Yang, Minghao, et al.
Published: (2025)
by: Yang, Minghao, et al.
Published: (2025)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
by: Gao, Yipeng, et al.
Published: (2023)
by: Gao, Yipeng, et al.
Published: (2023)
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
Chimera: Improving Generalist Model with Domain-Specific Experts
by: Peng, Tianshuo, et al.
Published: (2024)
by: Peng, Tianshuo, et al.
Published: (2024)
3D Scene Graph Guided Vision-Language Pre-training
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
by: Zhao, Chenhui, et al.
Published: (2025)
by: Zhao, Chenhui, et al.
Published: (2025)
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
by: Xue, Le, et al.
Published: (2023)
by: Xue, Le, et al.
Published: (2023)
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
by: Yamada, Ryousuke, et al.
Published: (2025)
by: Yamada, Ryousuke, et al.
Published: (2025)
Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward
by: Chen, Jianlong, et al.
Published: (2026)
by: Chen, Jianlong, et al.
Published: (2026)
Fully Sparse 3D Occupancy Prediction
by: Liu, Haisong, et al.
Published: (2023)
by: Liu, Haisong, et al.
Published: (2023)
RulePlanner: All-in-One Reinforcement Learner for Unifying Design Rules in 3D Floorplanning
by: Zhong, Ruizhe, et al.
Published: (2026)
by: Zhong, Ruizhe, et al.
Published: (2026)
Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
BePo: Dual Representation for 3D Occupancy Prediction
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors
by: Dong, Xin, et al.
Published: (2026)
by: Dong, Xin, et al.
Published: (2026)
Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
by: Zhang, Shaofeng, et al.
Published: (2025)
by: Zhang, Shaofeng, et al.
Published: (2025)
Position: Towards Implicit Prompt For Text-To-Image Models
by: Yang, Yue, et al.
Published: (2024)
by: Yang, Yue, et al.
Published: (2024)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
GaussianOcc: Fully Self-supervised and Efficient 3D Occupancy Estimation with Gaussian Splatting
by: Gan, Wanshui, et al.
Published: (2024)
by: Gan, Wanshui, et al.
Published: (2024)
OccupancyDETR: Using DETR for Mixed Dense-sparse 3D Occupancy Prediction
by: Jia, Yupeng, et al.
Published: (2023)
by: Jia, Yupeng, et al.
Published: (2023)
P3P: Pseudo-3D Pre-training for Scaling 3D Voxel-based Masked Autoencoders
by: Chen, Xuechao, et al.
Published: (2024)
by: Chen, Xuechao, et al.
Published: (2024)
STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction
by: Liao, Zhimin, et al.
Published: (2025)
by: Liao, Zhimin, et al.
Published: (2025)
Similar Items
-
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
by: Ye, Hancheng, et al.
Published: (2024) -
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
by: Xia, Renqiu, et al.
Published: (2024) -
ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
by: Zhang, Bo, et al.
Published: (2023) -
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
by: Xia, Renqiu, et al.
Published: (2024) -
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
by: Xia, Renqiu, et al.
Published: (2023)