Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Liwen, Yang, Weidong, Ma, Lipeng, Fei, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
by: Ma, Shuailei, et al.
Published: (2023)
by: Ma, Shuailei, et al.
Published: (2023)
3DMambaComplete: Exploring Structured State Space Model for Point Cloud Completion
by: Li, Yixuan, et al.
Published: (2024)
by: Li, Yixuan, et al.
Published: (2024)
MultiMAE Meets Earth Observation: Pre-training Multi-modal Multi-task Masked Autoencoders for Earth Observation Tasks
by: Sosa, Jose, et al.
Published: (2025)
by: Sosa, Jose, et al.
Published: (2025)
GS-PT: Exploiting 3D Gaussian Splatting for Comprehensive Point Cloud Understanding via Self-supervised Learning
by: Liu, Keyi, et al.
Published: (2024)
by: Liu, Keyi, et al.
Published: (2024)
MBPU: A Plug-and-Play State Space Model for Point Cloud Upsamping with Fast Point Rendering
by: Song, Jiayi, et al.
Published: (2024)
by: Song, Jiayi, et al.
Published: (2024)
Point Cloud Unsupervised Pre-training via 3D Gaussian Splatting
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts
by: Li, Weiran, et al.
Published: (2025)
by: Li, Weiran, et al.
Published: (2025)
Unified Multi-modal Diagnostic Framework with Reconstruction Pre-training and Heterogeneity-combat Tuning
by: Zhang, Yupei, et al.
Published: (2024)
by: Zhang, Yupei, et al.
Published: (2024)
Integrating Text and Image Pre-training for Multi-modal Algorithmic Reasoning
by: Zhang, Zijian, et al.
Published: (2024)
by: Zhang, Zijian, et al.
Published: (2024)
Adapting Multi-modal Large Language Model to Concept Drift From Pre-training Onwards
by: Yang, Xiaoyu, et al.
Published: (2024)
by: Yang, Xiaoyu, et al.
Published: (2024)
Robust Fine-tuning for Pre-trained 3D Point Cloud Models
by: Zhang, Zhibo, et al.
Published: (2024)
by: Zhang, Zhibo, et al.
Published: (2024)
Point Cloud as a Foreign Language for Multi-modal Large Language Model
by: Paul, Sneha, et al.
Published: (2026)
by: Paul, Sneha, et al.
Published: (2026)
Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
by: Shan, Ziyu, et al.
Published: (2024)
by: Shan, Ziyu, et al.
Published: (2024)
Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives
by: Luo, Sheng, et al.
Published: (2024)
by: Luo, Sheng, et al.
Published: (2024)
Pre-training Point Cloud Compact Model with Partial-aware Reconstruction
by: Zha, Yaohua, et al.
Published: (2024)
by: Zha, Yaohua, et al.
Published: (2024)
Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
by: Zhong, Xinliu, et al.
Published: (2025)
by: Zhong, Xinliu, et al.
Published: (2025)
Multi-modal Vision Pre-training for Medical Image Analysis
by: Rui, Shaohao, et al.
Published: (2024)
by: Rui, Shaohao, et al.
Published: (2024)
PIM: Physics-Informed Multi-task Pre-training for Improving Inertial Sensor-Based Human Activity Recognition
by: Nshimyimana, Dominique, et al.
Published: (2025)
by: Nshimyimana, Dominique, et al.
Published: (2025)
Multi-View Representation is What You Need for Point-Cloud Pre-Training
by: Yan, Siming, et al.
Published: (2023)
by: Yan, Siming, et al.
Published: (2023)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction
by: Gu, Lipeng, et al.
Published: (2024)
by: Gu, Lipeng, et al.
Published: (2024)
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
by: Xuan, Shiyu, et al.
Published: (2025)
by: Xuan, Shiyu, et al.
Published: (2025)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
by: Lin, Zhiwei, et al.
Published: (2022)
by: Lin, Zhiwei, et al.
Published: (2022)
PointGauss: Point Cloud-Guided Multi-Object Segmentation for Gaussian Splatting
by: Sun, Wentao, et al.
Published: (2025)
by: Sun, Wentao, et al.
Published: (2025)
InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
by: Luo, Zihao, et al.
Published: (2025)
by: Luo, Zihao, et al.
Published: (2025)
Explicitly Guided Information Interaction Network for Cross-modal Point Cloud Completion
by: Xu, Hang, et al.
Published: (2024)
by: Xu, Hang, et al.
Published: (2024)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
by: Walimbe, Soham, et al.
Published: (2025)
by: Walimbe, Soham, et al.
Published: (2025)
3DInAction: Understanding Human Actions in 3D Point Clouds
by: Ben-Shabat, Yizhak, et al.
Published: (2023)
by: Ben-Shabat, Yizhak, et al.
Published: (2023)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
by: Qiu, Xuerui, et al.
Published: (2026)
by: Qiu, Xuerui, et al.
Published: (2026)
Decision PCR: Decision version of the Point Cloud Registration task
by: Zhang, Yaojie, et al.
Published: (2025)
by: Zhang, Yaojie, et al.
Published: (2025)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
by: Miao, Yongzhu, et al.
Published: (2023)
by: Miao, Yongzhu, et al.
Published: (2023)
Mamba Learns in Context: Structure-Aware Domain Generalization for Multi-Task Point Cloud Understanding
by: Jiang, Jincen, et al.
Published: (2026)
by: Jiang, Jincen, et al.
Published: (2026)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
PCoTTA: Continual Test-Time Adaptation for Multi-Task Point Cloud Understanding
by: Jiang, Jincen, et al.
Published: (2024)
by: Jiang, Jincen, et al.
Published: (2024)
Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion
by: Hou, Hongye, et al.
Published: (2025)
by: Hou, Hongye, et al.
Published: (2025)
GeoAuxNet: Towards Universal 3D Representation Learning for Multi-sensor Point Clouds
by: Zhang, Shengjun, et al.
Published: (2024)
by: Zhang, Shengjun, et al.
Published: (2024)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
Similar Items
-
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
by: Ma, Shuailei, et al.
Published: (2023) -
3DMambaComplete: Exploring Structured State Space Model for Point Cloud Completion
by: Li, Yixuan, et al.
Published: (2024) -
MultiMAE Meets Earth Observation: Pre-training Multi-modal Multi-task Masked Autoencoders for Earth Observation Tasks
by: Sosa, Jose, et al.
Published: (2025) -
GS-PT: Exploiting 3D Gaussian Splatting for Comprehensive Point Cloud Understanding via Self-supervised Learning
by: Liu, Keyi, et al.
Published: (2024) -
MBPU: A Plug-and-Play State Space Model for Point Cloud Upsamping with Fast Point Rendering
by: Song, Jiayi, et al.
Published: (2024)