Delving into Spectral Clustering with Vision-Language Representations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Bo, Hu, Yuanwei, Liu, Bo, Chen, Ling, Lu, Jie, Fang, Zhen
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918389131247616
author Peng, Bo
Hu, Yuanwei
Liu, Bo
Chen, Ling
Lu, Jie
Fang, Zhen
author_facet Peng, Bo
Hu, Yuanwei
Liu, Bo
Chen, Ling
Lu, Jie
Fang, Zhen
contents Spectral clustering is known as a powerful technique in unsupervised data analysis. The vast majority of approaches to spectral clustering are driven by a single modality, leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of spectral clustering from a single-modal to a multi-modal regime. Particularly, we propose Neural Tangent Kernel Spectral Clustering that leverages cross-modal alignment in pre-trained vision-language models. By anchoring the neural tangent kernel with positive nouns, i.e., those semantically close to the images of interest, we arrive at formulating the affinity between images as a coupling of their visual proximity and semantic overlap. We show that this formulation amplifies within-cluster connections while suppressing spurious ones across clusters, hence encouraging block-diagonal structures. In addition, we present a regularized affinity diffusion mechanism that adaptively ensembles affinity matrices induced by different prompts. Extensive experiments on \textbf{16} benchmarks -- including classical, large-scale, fine-grained and domain-shifted datasets -- manifest that our method consistently outperforms the state-of-the-art by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09586
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Delving into Spectral Clustering with Vision-Language Representations
Peng, Bo
Hu, Yuanwei
Liu, Bo
Chen, Ling
Lu, Jie
Fang, Zhen
Computer Vision and Pattern Recognition
Spectral clustering is known as a powerful technique in unsupervised data analysis. The vast majority of approaches to spectral clustering are driven by a single modality, leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of spectral clustering from a single-modal to a multi-modal regime. Particularly, we propose Neural Tangent Kernel Spectral Clustering that leverages cross-modal alignment in pre-trained vision-language models. By anchoring the neural tangent kernel with positive nouns, i.e., those semantically close to the images of interest, we arrive at formulating the affinity between images as a coupling of their visual proximity and semantic overlap. We show that this formulation amplifies within-cluster connections while suppressing spurious ones across clusters, hence encouraging block-diagonal structures. In addition, we present a regularized affinity diffusion mechanism that adaptively ensembles affinity matrices induced by different prompts. Extensive experiments on \textbf{16} benchmarks -- including classical, large-scale, fine-grained and domain-shifted datasets -- manifest that our method consistently outperforms the state-of-the-art by a large margin.
title Delving into Spectral Clustering with Vision-Language Representations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.09586