Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Mei, Guofeng, Ren, Bin, Liu, Juan, Riz, Luigi, Huang, Xiaoshui, Zheng, Xu, Gong, Yongshun, Yang, Ming-Hsuan, Sebe, Nicu, Poiesi, Fabio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Geometrically-driven Aggregation for Zero-shot 3D Point Cloud Understanding
by: Mei, Guofeng, et al.
Published: (2023)
by: Mei, Guofeng, et al.
Published: (2023)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
by: Mei, Guofeng, et al.
Published: (2024)
by: Mei, Guofeng, et al.
Published: (2024)
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
by: Wang, Weijie, et al.
Published: (2023)
by: Wang, Weijie, et al.
Published: (2023)
PerLA: Perceptive 3D Language Assistant
by: Mei, Guofeng, et al.
Published: (2024)
by: Mei, Guofeng, et al.
Published: (2024)
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
by: Mei, Guofeng, et al.
Published: (2026)
by: Mei, Guofeng, et al.
Published: (2026)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
Fully-Geometric Cross-Attention for Point Cloud Registration
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
by: Qu, Wentao, et al.
Published: (2025)
by: Qu, Wentao, et al.
Published: (2025)
Novel class discovery meets foundation models for 3D semantic segmentation
by: Riz, Luigi, et al.
Published: (2023)
by: Riz, Luigi, et al.
Published: (2023)
3D Part Segmentation via Geometric Aggregation of 2D Visual Features
by: Garosi, Marco, et al.
Published: (2024)
by: Garosi, Marco, et al.
Published: (2024)
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
Diverse Teacher-Students for Deep Safe Semi-Supervised Learning under Class Mismatch
by: Wang, Qikai, et al.
Published: (2024)
by: Wang, Qikai, et al.
Published: (2024)
3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
by: Xue, Feng, et al.
Published: (2025)
by: Xue, Feng, et al.
Published: (2025)
Multimodal Fusion SLAM with Fourier Attention
by: Zhou, Youjie, et al.
Published: (2025)
by: Zhou, Youjie, et al.
Published: (2025)
Unsupervised Point Cloud Pre-Training via Contrasting and Clustering
by: Mei, Guofeng, et al.
Published: (2022)
by: Mei, Guofeng, et al.
Published: (2022)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025)
by: Zuo, Zhi, et al.
Published: (2025)
ShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
Fate of localization in coupled free chain and disordered chain
by: Lin, Xiaoshui, et al.
Published: (2023)
by: Lin, Xiaoshui, et al.
Published: (2023)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion
by: Qu, Wentao, et al.
Published: (2025)
by: Qu, Wentao, et al.
Published: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
Hierarchical Information Flow for Generalized Efficient Image Restoration
by: Li, Yawei, et al.
Published: (2024)
by: Li, Yawei, et al.
Published: (2024)
Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration
by: Li, Yawei, et al.
Published: (2025)
by: Li, Yawei, et al.
Published: (2025)
Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation
by: Li, Wenhao, et al.
Published: (2023)
by: Li, Wenhao, et al.
Published: (2023)
PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud Registration
by: Huang, Xiaoshui, et al.
Published: (2025)
by: Huang, Xiaoshui, et al.
Published: (2025)
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
by: Liu, Dingning, et al.
Published: (2024)
by: Liu, Dingning, et al.
Published: (2024)
An Efficient Learning-based Solver Comparable to Metaheuristics for the Capacitated Arc Routing Problem
by: Guo, Runze, et al.
Published: (2024)
by: Guo, Runze, et al.
Published: (2024)
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
Analysis of Indistinguishable Trajectories of a Nonholonomic Vehicle Subject to Range Measurements
by: Riz, Francesco, et al.
Published: (2022)
by: Riz, Francesco, et al.
Published: (2022)
Action-guided generation of 3D functionality segmentation data
by: Corsetti, Jaime, et al.
Published: (2025)
by: Corsetti, Jaime, et al.
Published: (2025)
Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving
by: Sun, Jiangxin, et al.
Published: (2026)
by: Sun, Jiangxin, et al.
Published: (2026)
On the different Floquet Hamiltonians in a periodic-driven Bose-Josephson junction
by: Lin, Xiaoshui, et al.
Published: (2023)
by: Lin, Xiaoshui, et al.
Published: (2023)
Landau-Zener-Stückelberg interference in edge state pumping
by: Liu, Y., et al.
Published: (2024)
by: Liu, Y., et al.
Published: (2024)
Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
by: Lv, Chonghua, et al.
Published: (2026)
by: Lv, Chonghua, et al.
Published: (2026)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
by: Song, Yue, et al.
Published: (2023)
by: Song, Yue, et al.
Published: (2023)
Similar Items
-
Geometrically-driven Aggregation for Zero-shot 3D Point Cloud Understanding
by: Mei, Guofeng, et al.
Published: (2023) -
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025) -
Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
by: Mei, Guofeng, et al.
Published: (2024) -
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
by: Wang, Weijie, et al.
Published: (2023) -
PerLA: Perceptive 3D Language Assistant
by: Mei, Guofeng, et al.
Published: (2024)