Label Propagation for Zero-shot Classification with Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stojnić, Vladan, Kalantidis, Yannis, Tolias, Giorgos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909161984360448
author Stojnić, Vladan
Kalantidis, Yannis
Tolias, Giorgos
author_facet Stojnić, Vladan
Kalantidis, Yannis
Tolias, Giorgos
contents Vision-Language Models (VLMs) have demonstrated impressive performance on zero-shot classification, i.e. classification when provided merely with a list of class names. In this paper, we tackle the case of zero-shot classification in the presence of unlabeled data. We leverage the graph structure of the unlabeled data and introduce ZLaP, a method based on label propagation (LP) that utilizes geodesic distances for classification. We tailor LP to graphs containing both text and image features and further propose an efficient method for performing inductive inference based on a dual solution and a sparsification step. We perform extensive experiments to evaluate the effectiveness of our method on 14 common datasets and show that ZLaP outperforms the latest related works. Code: https://github.com/vladan-stojnic/ZLaP
format Preprint
id arxiv_https___arxiv_org_abs_2404_04072
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Label Propagation for Zero-shot Classification with Vision-Language Models
Stojnić, Vladan
Kalantidis, Yannis
Tolias, Giorgos
Computer Vision and Pattern Recognition
Machine Learning
Vision-Language Models (VLMs) have demonstrated impressive performance on zero-shot classification, i.e. classification when provided merely with a list of class names. In this paper, we tackle the case of zero-shot classification in the presence of unlabeled data. We leverage the graph structure of the unlabeled data and introduce ZLaP, a method based on label propagation (LP) that utilizes geodesic distances for classification. We tailor LP to graphs containing both text and image features and further propose an efficient method for performing inductive inference based on a dual solution and a sparsification step. We perform extensive experiments to evaluate the effectiveness of our method on 14 common datasets and show that ZLaP outperforms the latest related works. Code: https://github.com/vladan-stojnic/ZLaP
title Label Propagation for Zero-shot Classification with Vision-Language Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2404.04072