UniM$^2$AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Jian, Huang, Tianyu, Yang, Guanglei, Guo, Zhenhua, Luo, Tao, Feng, Chun-Mei, Zuo, Wangmeng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models
by: Xu, Wan, et al.
Published: (2023)
by: Xu, Wan, et al.
Published: (2023)
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
by: Zheng, Mi, et al.
Published: (2025)
by: Zheng, Mi, et al.
Published: (2025)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
by: Li, Yuanze, et al.
Published: (2024)
by: Li, Yuanze, et al.
Published: (2024)
S2AM3D: Scale-controllable Part Segmentation of 3D Point Clouds
by: Su, Han, et al.
Published: (2025)
by: Su, Han, et al.
Published: (2025)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
CMViM: Contrastive Masked Vim Autoencoder for 3D Multi-modal Representation Learning for AD classification
by: Yang, Guangqian, et al.
Published: (2024)
by: Yang, Guangqian, et al.
Published: (2024)
V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving
by: Luo, Xuewen, et al.
Published: (2025)
by: Luo, Xuewen, et al.
Published: (2025)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
by: Zou, Jialv, et al.
Published: (2024)
by: Zou, Jialv, et al.
Published: (2024)
Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer
by: Wan, Yecong, et al.
Published: (2025)
by: Wan, Yecong, et al.
Published: (2025)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
by: Min, Chen, et al.
Published: (2023)
by: Min, Chen, et al.
Published: (2023)
FedSmoothLoRA: Toward Smoother and Faster Convergence in Federated Low-Rank Adaptation
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
by: Li, Yingwei, et al.
Published: (2026)
by: Li, Yingwei, et al.
Published: (2026)
CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving
by: Chen, Runjian, et al.
Published: (2022)
by: Chen, Runjian, et al.
Published: (2022)
UniVision: A Unified Framework for Vision-Centric 3D Perception
by: Hong, Yu, et al.
Published: (2024)
by: Hong, Yu, et al.
Published: (2024)
OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving
by: Wang, Guoqing, et al.
Published: (2024)
by: Wang, Guoqing, et al.
Published: (2024)
Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
by: Yang, Feng, et al.
Published: (2025)
by: Yang, Feng, et al.
Published: (2025)
ConSept: Continual Semantic Segmentation via Adapter-based Vision Transformer
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
MetricDepth: Enhancing Monocular Depth Estimation with Deep Metric Learning
by: Liu, Chunpu, et al.
Published: (2024)
by: Liu, Chunpu, et al.
Published: (2024)
UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
by: Xiong, Zhexiao, et al.
Published: (2026)
by: Xiong, Zhexiao, et al.
Published: (2026)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
by: Goswami, Prajnan, et al.
Published: (2026)
by: Goswami, Prajnan, et al.
Published: (2026)
ZFusion: An Effective Fuser of Camera and 4D Radar for 3D Object Perception in Autonomous Driving
by: Yang, Sheng, et al.
Published: (2025)
by: Yang, Sheng, et al.
Published: (2025)
TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation
by: Jin, Bu, et al.
Published: (2025)
by: Jin, Bu, et al.
Published: (2025)
DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
Toward Unified Multimodal Representation Learning for Autonomous Driving
by: Tao, Ximeng, et al.
Published: (2026)
by: Tao, Ximeng, et al.
Published: (2026)
UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
by: Lu, Hao, et al.
Published: (2025)
by: Lu, Hao, et al.
Published: (2025)
RadarMP: Motion Perception for 4D mmWave Radar in Autonomous Driving
by: Cheng, Ruiqi, et al.
Published: (2025)
by: Cheng, Ruiqi, et al.
Published: (2025)
Multi-modal Relation Distillation for Unified 3D Representation Learning
by: Wang, Huiqun, et al.
Published: (2024)
by: Wang, Huiqun, et al.
Published: (2024)
ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2025)
by: Ma, Yunsheng, et al.
Published: (2025)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
Pseudo Labelling for Enhanced Masked Autoencoders
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving
by: Wang, Xiyang, et al.
Published: (2024)
by: Wang, Xiyang, et al.
Published: (2024)
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
by: Gao, Sensen, et al.
Published: (2026)
by: Gao, Sensen, et al.
Published: (2026)
S2Gaussian: Sparse-View Super-Resolution 3D Gaussian Splatting
by: Wan, Yecong, et al.
Published: (2025)
by: Wan, Yecong, et al.
Published: (2025)
GaussianPretrain: A Simple Unified 3D Gaussian Representation for Visual Pre-training in Autonomous Driving
by: Xu, Shaoqing, et al.
Published: (2024)
by: Xu, Shaoqing, et al.
Published: (2024)
A Comprehensive Survey on 3D Content Generation
by: Liu, Jian, et al.
Published: (2024)
by: Liu, Jian, et al.
Published: (2024)
UniFuture: A 4D Driving World Model for Future Generation and Perception
by: Liang, Dingkang, et al.
Published: (2025)
by: Liang, Dingkang, et al.
Published: (2025)
Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception
by: Bhowmick, Arka, et al.
Published: (2026)
by: Bhowmick, Arka, et al.
Published: (2026)
Similar Items
-
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
by: He, Qingdong, et al.
Published: (2024) -
FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models
by: Xu, Wan, et al.
Published: (2023) -
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
by: Zheng, Mi, et al.
Published: (2025) -
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026) -
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
by: Li, Yuanze, et al.
Published: (2024)