Saved in:
Bibliographic Details
Main Authors: Xiao, Qinfeng, Mei, Guofeng, Liu, Qilong, Yi, Chenyuan, Poiesi, Fabio, Zhang, Jian, Yang, Bo, Kit-lun, Yick
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.07652
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914378701340672
author Xiao, Qinfeng
Mei, Guofeng
Liu, Qilong
Yi, Chenyuan
Poiesi, Fabio
Zhang, Jian
Yang, Bo
Kit-lun, Yick
author_facet Xiao, Qinfeng
Mei, Guofeng
Liu, Qilong
Yi, Chenyuan
Poiesi, Fabio
Zhang, Jian
Yang, Bo
Kit-lun, Yick
contents Establishing dense correspondence across 3D shapes is crucial for fundamental downstream tasks, including texture transfer, shape interpolation, and robotic manipulation. However, learning these mappings without manual supervision remains a formidable challenge, particularly under severe non-isometric deformations and in inter-class settings where geometric cues are ambiguous. Conventional functional map methods, while elegant, typically struggle in these regimes due to their reliance on isometry. To address this, we present GLASS, a framework that bridges the gap by integrating geometric spectral analysis with rich semantic priors from vision-language foundation models. GLASS introduces three key innovations: (i) a view-consistent strategy that enables robust multi-view visual feature extraction from powerful vision foundation models; (ii) the injection of language embeddings into vertex descriptors via zero-shot 3D segmentation, capturing high-level part semantics; and (iii) a graph-assisted contrastive loss that enforces structural consistency between regions (e.g., source's head'' $\leftrightarrow$ target's head'') by leveraging geodesic and topological relationships between regions. This design allows GLASS to learn globally coherent and semantically consistent maps without ground-truth supervision. Extensive experiments demonstrate that GLASS achieves state-of-the-art performance across all regimes, maintaining high accuracy on standard near-isometric tasks while significantly advancing performance in challenging settings. Specifically, it achieves average geodesic errors of 0.21, 4.5, and 5.6 on the inter-class benchmark SNIS and non-isometric benchmarks SMAL and TOPKIDS, reducing errors from URSSM baselines of 0.49, 6.0, and 8.9 by 57%, 25%, and 37%, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2603_07652
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GLASS: Graph and Vision-Language Assisted Semantic Shape Correspondence
Xiao, Qinfeng
Mei, Guofeng
Liu, Qilong
Yi, Chenyuan
Poiesi, Fabio
Zhang, Jian
Yang, Bo
Kit-lun, Yick
Computer Vision and Pattern Recognition
Establishing dense correspondence across 3D shapes is crucial for fundamental downstream tasks, including texture transfer, shape interpolation, and robotic manipulation. However, learning these mappings without manual supervision remains a formidable challenge, particularly under severe non-isometric deformations and in inter-class settings where geometric cues are ambiguous. Conventional functional map methods, while elegant, typically struggle in these regimes due to their reliance on isometry. To address this, we present GLASS, a framework that bridges the gap by integrating geometric spectral analysis with rich semantic priors from vision-language foundation models. GLASS introduces three key innovations: (i) a view-consistent strategy that enables robust multi-view visual feature extraction from powerful vision foundation models; (ii) the injection of language embeddings into vertex descriptors via zero-shot 3D segmentation, capturing high-level part semantics; and (iii) a graph-assisted contrastive loss that enforces structural consistency between regions (e.g., source's head'' $\leftrightarrow$ target's head'') by leveraging geodesic and topological relationships between regions. This design allows GLASS to learn globally coherent and semantically consistent maps without ground-truth supervision. Extensive experiments demonstrate that GLASS achieves state-of-the-art performance across all regimes, maintaining high accuracy on standard near-isometric tasks while significantly advancing performance in challenging settings. Specifically, it achieves average geodesic errors of 0.21, 4.5, and 5.6 on the inter-class benchmark SNIS and non-isometric benchmarks SMAL and TOPKIDS, reducing errors from URSSM baselines of 0.49, 6.0, and 8.9 by 57%, 25%, and 37%, respectively.
title GLASS: Graph and Vision-Language Assisted Semantic Shape Correspondence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.07652