Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mei, Guofeng, Ren, Bin, Liu, Juan, Riz, Luigi, Huang, Xiaoshui, Zheng, Xu, Gong, Yongshun, Yang, Ming-Hsuan, Sebe, Nicu, Poiesi, Fabio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915303573684224
author Mei, Guofeng
Ren, Bin
Liu, Juan
Riz, Luigi
Huang, Xiaoshui
Zheng, Xu
Gong, Yongshun
Yang, Ming-Hsuan
Sebe, Nicu
Poiesi, Fabio
author_facet Mei, Guofeng
Ren, Bin
Liu, Juan
Riz, Luigi
Huang, Xiaoshui
Zheng, Xu
Gong, Yongshun
Yang, Ming-Hsuan
Sebe, Nicu
Poiesi, Fabio
contents Vision-language models like CLIP can offer a promising foundation for 3D scene understanding when extended with 3D tokenizers. However, standard approaches, such as k-nearest neighbor or radius-based tokenization, struggle with cross-domain generalization due to sensitivity to dataset-specific spatial scales. We present a universal 3D tokenizer designed for scale-invariant representation learning with a frozen CLIP backbone. We show that combining superpoint-based grouping with coordinate scale normalization consistently outperforms conventional methods through extensive experimental analysis. Specifically, we introduce S4Token, a tokenization pipeline that produces semantically-informed tokens regardless of scene scale. Our tokenizer is trained without annotations using masked point modeling and clustering-based objectives, along with cross-modal distillation to align 3D tokens with 2D multi-view image features. For dense prediction tasks, we propose a superpoint-level feature propagation module to recover point-level detail from sparse tokens.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
Mei, Guofeng
Ren, Bin
Liu, Juan
Riz, Luigi
Huang, Xiaoshui
Zheng, Xu
Gong, Yongshun
Yang, Ming-Hsuan
Sebe, Nicu
Poiesi, Fabio
Computer Vision and Pattern Recognition
Vision-language models like CLIP can offer a promising foundation for 3D scene understanding when extended with 3D tokenizers. However, standard approaches, such as k-nearest neighbor or radius-based tokenization, struggle with cross-domain generalization due to sensitivity to dataset-specific spatial scales. We present a universal 3D tokenizer designed for scale-invariant representation learning with a frozen CLIP backbone. We show that combining superpoint-based grouping with coordinate scale normalization consistently outperforms conventional methods through extensive experimental analysis. Specifically, we introduce S4Token, a tokenization pipeline that produces semantically-informed tokens regardless of scene scale. Our tokenizer is trained without annotations using masked point modeling and clustering-based objectives, along with cross-modal distillation to align 3D tokens with 2D multi-view image features. For dense prediction tasks, we propose a superpoint-level feature propagation module to recover point-level detail from sparse tokens.
title Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18819