MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Zou, Jialv, Liao, Bencheng, Zhang, Qian, Liu, Wenyu, Wang, Xinggang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
di: Zou, Jialv, et al.
Pubblicazione: (2025)
di: Zou, Jialv, et al.
Pubblicazione: (2025)
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
di: Zou, Jialv, et al.
Pubblicazione: (2025)
di: Zou, Jialv, et al.
Pubblicazione: (2025)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
di: Jiang, Bo, et al.
Pubblicazione: (2024)
di: Jiang, Bo, et al.
Pubblicazione: (2024)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
di: Li, Yingyue, et al.
Pubblicazione: (2025)
di: Li, Yingyue, et al.
Pubblicazione: (2025)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
di: Jiang, Bo, et al.
Pubblicazione: (2024)
di: Jiang, Bo, et al.
Pubblicazione: (2024)
ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
di: Zheng, Zhiyu, et al.
Pubblicazione: (2025)
di: Zheng, Zhiyu, et al.
Pubblicazione: (2025)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
di: Jiang, Bo, et al.
Pubblicazione: (2025)
di: Jiang, Bo, et al.
Pubblicazione: (2025)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
di: Zeng, Lunbin, et al.
Pubblicazione: (2025)
di: Zeng, Lunbin, et al.
Pubblicazione: (2025)
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
di: Liao, Bencheng, et al.
Pubblicazione: (2024)
di: Liao, Bencheng, et al.
Pubblicazione: (2024)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
di: Liao, Bencheng, et al.
Pubblicazione: (2025)
di: Liao, Bencheng, et al.
Pubblicazione: (2025)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
di: Liao, Bencheng, et al.
Pubblicazione: (2024)
di: Liao, Bencheng, et al.
Pubblicazione: (2024)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
di: Li, Yongkang, et al.
Pubblicazione: (2026)
di: Li, Yongkang, et al.
Pubblicazione: (2026)
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
Polar Parametrization for Vision-based Surround-View 3D Detection
di: Chen, Shaoyu, et al.
Pubblicazione: (2022)
di: Chen, Shaoyu, et al.
Pubblicazione: (2022)
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
di: Li, Yongkang, et al.
Pubblicazione: (2024)
di: Li, Yongkang, et al.
Pubblicazione: (2024)
MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
di: Zhang, Wenrui, et al.
Pubblicazione: (2025)
di: Zhang, Wenrui, et al.
Pubblicazione: (2025)
RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
di: Gao, Hao, et al.
Pubblicazione: (2025)
di: Gao, Hao, et al.
Pubblicazione: (2025)
Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning
di: Song, Yuehao, et al.
Pubblicazione: (2026)
di: Song, Yuehao, et al.
Pubblicazione: (2026)
STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
di: Deng, Yunze, et al.
Pubblicazione: (2025)
di: Deng, Yunze, et al.
Pubblicazione: (2025)
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
di: Zou, Ya, et al.
Pubblicazione: (2025)
di: Zou, Ya, et al.
Pubblicazione: (2025)
GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
di: Cheng, Tianheng, et al.
Pubblicazione: (2026)
di: Cheng, Tianheng, et al.
Pubblicazione: (2026)
CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding
di: Zhang, Mingming, et al.
Pubblicazione: (2023)
di: Zhang, Mingming, et al.
Pubblicazione: (2023)
SemanticMIM: Marring Masked Image Modeling with Semantics Compression for General Visual Representation
di: Yuan, Yike, et al.
Pubblicazione: (2024)
di: Yuan, Yike, et al.
Pubblicazione: (2024)
Fully Unified Motion Planning for End-to-End Autonomous Driving
di: Liu, Lin, et al.
Pubblicazione: (2025)
di: Liu, Lin, et al.
Pubblicazione: (2025)
UniM$^2$AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
di: Zou, Jian, et al.
Pubblicazione: (2023)
di: Zou, Jian, et al.
Pubblicazione: (2023)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
di: Hu, Bin, et al.
Pubblicazione: (2024)
di: Hu, Bin, et al.
Pubblicazione: (2024)
Causality-inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
di: Xiong, Haijun, et al.
Pubblicazione: (2024)
di: Xiong, Haijun, et al.
Pubblicazione: (2024)
PiLaMIM: Toward Richer Visual Representations by Integrating Pixel and Latent Masked Image Modeling
di: Lee, Junmyeong, et al.
Pubblicazione: (2025)
di: Lee, Junmyeong, et al.
Pubblicazione: (2025)
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
di: Yao, Jingfeng, et al.
Pubblicazione: (2023)
di: Yao, Jingfeng, et al.
Pubblicazione: (2023)
VasoMIM: Vascular Anatomy-Aware Masked Image Modeling for Vessel Segmentation
di: Huang, De-Xing, et al.
Pubblicazione: (2025)
di: Huang, De-Xing, et al.
Pubblicazione: (2025)
Occupancy as Set of Points
di: Shi, Yiang, et al.
Pubblicazione: (2024)
di: Shi, Yiang, et al.
Pubblicazione: (2024)
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
di: Xia, Tianze, et al.
Pubblicazione: (2025)
di: Xia, Tianze, et al.
Pubblicazione: (2025)
PR-MIM: Delving Deeper into Partial Reconstruction in Masked Image Modeling
di: Li, Zhong-Yu, et al.
Pubblicazione: (2024)
di: Li, Zhong-Yu, et al.
Pubblicazione: (2024)
Adaptive Fusion of Single-View and Multi-View Depth for Autonomous Driving
di: Cheng, JunDa, et al.
Pubblicazione: (2024)
di: Cheng, JunDa, et al.
Pubblicazione: (2024)
AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation
di: Zhu, Lei, et al.
Pubblicazione: (2025)
di: Zhu, Lei, et al.
Pubblicazione: (2025)
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
di: Zou, Jialv, et al.
Pubblicazione: (2025) -
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
di: Zou, Jialv, et al.
Pubblicazione: (2025) -
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
di: Zhu, Lianghui, et al.
Pubblicazione: (2024) -
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
di: Jiang, Bo, et al.
Pubblicazione: (2024) -
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
di: Li, Yingyue, et al.
Pubblicazione: (2025)