CLIP's Visual Embedding Projector is a Few-shot Cornucopia

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fahes, Mohammad, Vu, Tuan-Hung, Bursuc, Andrei, Pérez, Patrick, de Charette, Raoul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908786965348352
author Fahes, Mohammad
Vu, Tuan-Hung
Bursuc, Andrei
Pérez, Patrick
de Charette, Raoul
author_facet Fahes, Mohammad
Vu, Tuan-Hung
Bursuc, Andrei
Pérez, Patrick
de Charette, Raoul
contents We introduce ProLIP, a simple and architecture-agnostic method for adapting contrastively pretrained vision-language models, such as CLIP, to few-shot classification. ProLIP fine-tunes the vision encoder's projection matrix with Frobenius norm regularization on its deviation from the pretrained weights. It achieves state-of-the-art performance on 11 few-shot classification benchmarks under both ``few-shot validation'' and ``validation-free'' settings. Moreover, by rethinking the non-linear CLIP-Adapter through ProLIP's lens, we design a Regularized Linear Adapter (RLA) that performs better, requires no hyperparameter tuning, is less sensitive to learning rate values, and offers an alternative to ProLIP in black-box scenarios where model weights are inaccessible. Beyond few-shot classification, ProLIP excels in cross-dataset transfer, domain generalization, base-to-new class generalization, and test-time adaptation--where it outperforms prompt tuning while being an order of magnitude faster to train. Code is available at https://github.com/astra-vision/ProLIP .
format Preprint
id arxiv_https___arxiv_org_abs_2410_05270
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIP's Visual Embedding Projector is a Few-shot Cornucopia
Fahes, Mohammad
Vu, Tuan-Hung
Bursuc, Andrei
Pérez, Patrick
de Charette, Raoul
Computer Vision and Pattern Recognition
We introduce ProLIP, a simple and architecture-agnostic method for adapting contrastively pretrained vision-language models, such as CLIP, to few-shot classification. ProLIP fine-tunes the vision encoder's projection matrix with Frobenius norm regularization on its deviation from the pretrained weights. It achieves state-of-the-art performance on 11 few-shot classification benchmarks under both ``few-shot validation'' and ``validation-free'' settings. Moreover, by rethinking the non-linear CLIP-Adapter through ProLIP's lens, we design a Regularized Linear Adapter (RLA) that performs better, requires no hyperparameter tuning, is less sensitive to learning rate values, and offers an alternative to ProLIP in black-box scenarios where model weights are inaccessible. Beyond few-shot classification, ProLIP excels in cross-dataset transfer, domain generalization, base-to-new class generalization, and test-time adaptation--where it outperforms prompt tuning while being an order of magnitude faster to train. Code is available at https://github.com/astra-vision/ProLIP .
title CLIP's Visual Embedding Projector is a Few-shot Cornucopia
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.05270