VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Su, Yuetong, Wei, Baoguo, Wang, Xinyu, Li, Xu, Li, Lixin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914194184470528
author Su, Yuetong
Wei, Baoguo
Wang, Xinyu
Li, Xu
Li, Lixin
author_facet Su, Yuetong
Wei, Baoguo
Wang, Xinyu
Li, Xu
Li, Lixin
contents Novel Class Discovery aims to utilise prior knowledge of known classes to classify and discover unknown classes from unlabelled data. Existing NCD methods for images primarily rely on visual features, which suffer from limitations such as insufficient feature discriminability and the long-tail distribution of data. We propose LLM-NCD, a multimodal framework that breaks this bottleneck by fusing visual-textual semantics and prototype guided clustering. Our key innovation lies in modelling cluster centres and semantic prototypes of known classes by jointly optimising known class image and text features, and a dualphase discovery mechanism that dynamically separates known or novel samples via semantic affinity thresholds and adaptive clustering. Experiments on the CIFAR-100 dataset show that compared to the current methods, this method achieves up to 25.3% improvement in accuracy for unknown classes. Notably, our method shows unique resilience to long tail distributions, a first in NCD literature.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
Su, Yuetong
Wei, Baoguo
Wang, Xinyu
Li, Xu
Li, Lixin
Computer Vision and Pattern Recognition
68T45
I.2.10; I.4.8; I.5
Novel Class Discovery aims to utilise prior knowledge of known classes to classify and discover unknown classes from unlabelled data. Existing NCD methods for images primarily rely on visual features, which suffer from limitations such as insufficient feature discriminability and the long-tail distribution of data. We propose LLM-NCD, a multimodal framework that breaks this bottleneck by fusing visual-textual semantics and prototype guided clustering. Our key innovation lies in modelling cluster centres and semantic prototypes of known classes by jointly optimising known class image and text features, and a dualphase discovery mechanism that dynamically separates known or novel samples via semantic affinity thresholds and adaptive clustering. Experiments on the CIFAR-100 dataset show that compared to the current methods, this method achieves up to 25.3% improvement in accuracy for unknown classes. Notably, our method shows unique resilience to long tail distributions, a first in NCD literature.
title VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
topic Computer Vision and Pattern Recognition
68T45
I.2.10; I.4.8; I.5
url https://arxiv.org/abs/2512.10262