Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive Setting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kravets, Alexey, Chen, Da, Namboodiri, Vinay P.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918105620414464
author Kravets, Alexey
Chen, Da
Namboodiri, Vinay P.
author_facet Kravets, Alexey
Chen, Da
Namboodiri, Vinay P.
contents CLIP is a foundational model with transferable classification performance in the few-shot setting. Several methods have shown improved performance of CLIP using few-shot examples. However, so far, all these techniques have been benchmarked using standard few-shot datasets. We argue that this mode of evaluation does not provide a true indication of the inductive generalization ability using few-shot examples. As most datasets have been seen by the CLIP model, the resultant setting can be termed as partially transductive. To solve this, we propose a pipeline that uses an unlearning technique to obtain true inductive baselines. In this new inductive setting, the methods show a significant drop in performance (-55% on average among 13 baselines with multiple datasets). We validate the unlearning technique using oracle baselines. An improved few-shot classification technique is proposed that consistently obtains state-of-the-art performance over 13 other recent baseline methods on a comprehensive analysis with 5880 experiments - varying the datasets, differing number of few-shot examples, unlearning setting, and with different seeds. Thus, we identify the issue with the evaluation of CLIP-based few-shot classification, provide a solution using unlearning, propose new benchmarks, and provide an improved method.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20834
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive Setting
Kravets, Alexey
Chen, Da
Namboodiri, Vinay P.
Computer Vision and Pattern Recognition
CLIP is a foundational model with transferable classification performance in the few-shot setting. Several methods have shown improved performance of CLIP using few-shot examples. However, so far, all these techniques have been benchmarked using standard few-shot datasets. We argue that this mode of evaluation does not provide a true indication of the inductive generalization ability using few-shot examples. As most datasets have been seen by the CLIP model, the resultant setting can be termed as partially transductive. To solve this, we propose a pipeline that uses an unlearning technique to obtain true inductive baselines. In this new inductive setting, the methods show a significant drop in performance (-55% on average among 13 baselines with multiple datasets). We validate the unlearning technique using oracle baselines. An improved few-shot classification technique is proposed that consistently obtains state-of-the-art performance over 13 other recent baseline methods on a comprehensive analysis with 5880 experiments - varying the datasets, differing number of few-shot examples, unlearning setting, and with different seeds. Thus, we identify the issue with the evaluation of CLIP-based few-shot classification, provide a solution using unlearning, propose new benchmarks, and provide an improved method.
title Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive Setting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.20834