Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yujia, Wu, Xiaoyang, Lao, Yixing, Wang, Chengyao, Tian, Zhuotao, Wang, Naiyan, Zhao, Hengshuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911475127287808
author Zhang, Yujia
Wu, Xiaoyang
Lao, Yixing
Wang, Chengyao
Tian, Zhuotao
Wang, Naiyan
Zhao, Hengshuang
author_facet Zhang, Yujia
Wu, Xiaoyang
Lao, Yixing
Wang, Chengyao
Tian, Zhuotao
Wang, Naiyan
Zhao, Hengshuang
contents Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-distillation with 2D-3D cross-modal joint embedding. Despite its simplicity, Concerto learns more coherent and informative spatial features, as demonstrated by zero-shot visualizations. It outperforms both standalone SOTA 2D and 3D self-supervised models by 14.2% and 4.8%, respectively, as well as their feature concatenation, in linear probing for 3D scene perception. With full fine-tuning, Concerto sets new SOTA results across multiple scene understanding benchmarks (e.g., 80.7% mIoU on ScanNet). We further present a variant of Concerto tailored for video-lifted point cloud spatial understanding, and a translator that linearly projects Concerto representations into CLIP's language space, enabling open-world perception. These results highlight that Concerto emerges spatial representations with superior fine-grained geometric and semantic consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23607
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Zhang, Yujia
Wu, Xiaoyang
Lao, Yixing
Wang, Chengyao
Tian, Zhuotao
Wang, Naiyan
Zhao, Hengshuang
Computer Vision and Pattern Recognition
Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-distillation with 2D-3D cross-modal joint embedding. Despite its simplicity, Concerto learns more coherent and informative spatial features, as demonstrated by zero-shot visualizations. It outperforms both standalone SOTA 2D and 3D self-supervised models by 14.2% and 4.8%, respectively, as well as their feature concatenation, in linear probing for 3D scene perception. With full fine-tuning, Concerto sets new SOTA results across multiple scene understanding benchmarks (e.g., 80.7% mIoU on ScanNet). We further present a variant of Concerto tailored for video-lifted point cloud spatial understanding, and a translator that linearly projects Concerto representations into CLIP's language space, enabling open-world perception. These results highlight that Concerto emerges spatial representations with superior fine-grained geometric and semantic consistency.
title Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.23607