Multimodal CLIP Inference for Meta-Few-Shot Image Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ferragu, Constance, Chagniot, Philomene, Coyette, Vincent
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910452606304256
author Ferragu, Constance
Chagniot, Philomene
Coyette, Vincent
author_facet Ferragu, Constance
Chagniot, Philomene
Coyette, Vincent
contents In recent literature, few-shot classification has predominantly been defined by the N-way k-shot meta-learning problem. Models designed for this purpose are usually trained to excel on standard benchmarks following a restricted setup, excluding the use of external data. Given the recent advancements in large language and vision models, a question naturally arises: can these models directly perform well on meta-few-shot learning benchmarks? Multimodal foundation models like CLIP, which learn a joint (image, text) embedding, are of particular interest. Indeed, multimodal training has proven to enhance model robustness, especially regarding ambiguities, a limitation frequently observed in the few-shot setup. This study demonstrates that combining modalities from CLIP's text and image encoders outperforms state-of-the-art meta-few-shot learners on widely adopted benchmarks, all without additional training. Our results confirm the potential and robustness of multimodal foundation models like CLIP and serve as a baseline for existing and future approaches leveraging such models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10954
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal CLIP Inference for Meta-Few-Shot Image Classification
Ferragu, Constance
Chagniot, Philomene
Coyette, Vincent
Computer Vision and Pattern Recognition
In recent literature, few-shot classification has predominantly been defined by the N-way k-shot meta-learning problem. Models designed for this purpose are usually trained to excel on standard benchmarks following a restricted setup, excluding the use of external data. Given the recent advancements in large language and vision models, a question naturally arises: can these models directly perform well on meta-few-shot learning benchmarks? Multimodal foundation models like CLIP, which learn a joint (image, text) embedding, are of particular interest. Indeed, multimodal training has proven to enhance model robustness, especially regarding ambiguities, a limitation frequently observed in the few-shot setup. This study demonstrates that combining modalities from CLIP's text and image encoders outperforms state-of-the-art meta-few-shot learners on widely adopted benchmarks, all without additional training. Our results confirm the potential and robustness of multimodal foundation models like CLIP and serve as a baseline for existing and future approaches leveraging such models.
title Multimodal CLIP Inference for Meta-Few-Shot Image Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.10954