LLM meets Vision-Language Models for Zero-Shot One-Class Classification

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bendou, Yassir, Lioi, Giulia, Pasdeloup, Bastien, Mauch, Lukas, Hacene, Ghouthi Boukli, Cardinaux, Fabien, Gripon, Vincent
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917675734663168
author Bendou, Yassir
Lioi, Giulia
Pasdeloup, Bastien
Mauch, Lukas
Hacene, Ghouthi Boukli
Cardinaux, Fabien
Gripon, Vincent
author_facet Bendou, Yassir
Lioi, Giulia
Pasdeloup, Bastien
Mauch, Lukas
Hacene, Ghouthi Boukli
Cardinaux, Fabien
Gripon, Vincent
contents We consider the problem of zero-shot one-class visual classification, extending traditional one-class classification to scenarios where only the label of the target class is available. This method aims to discriminate between positive and negative query samples without requiring examples from the target class. We propose a two-step solution that first queries large language models for visually confusing objects and then relies on vision-language pre-trained models (e.g., CLIP) to perform classification. By adapting large-scale vision benchmarks, we demonstrate the ability of the proposed method to outperform adapted off-the-shelf alternatives in this setting. Namely, we propose a realistic benchmark where negative query samples are drawn from the same original dataset as positive ones, including a granularity-controlled version of iNaturalist, where negative samples are at a fixed distance in the taxonomy tree from the positive ones. To our knowledge, we are the first to demonstrate the ability to discriminate a single category from other semantically related ones using only its label.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00675
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLM meets Vision-Language Models for Zero-Shot One-Class Classification
Bendou, Yassir
Lioi, Giulia
Pasdeloup, Bastien
Mauch, Lukas
Hacene, Ghouthi Boukli
Cardinaux, Fabien
Gripon, Vincent
Computer Vision and Pattern Recognition
Artificial Intelligence
We consider the problem of zero-shot one-class visual classification, extending traditional one-class classification to scenarios where only the label of the target class is available. This method aims to discriminate between positive and negative query samples without requiring examples from the target class. We propose a two-step solution that first queries large language models for visually confusing objects and then relies on vision-language pre-trained models (e.g., CLIP) to perform classification. By adapting large-scale vision benchmarks, we demonstrate the ability of the proposed method to outperform adapted off-the-shelf alternatives in this setting. Namely, we propose a realistic benchmark where negative query samples are drawn from the same original dataset as positive ones, including a granularity-controlled version of iNaturalist, where negative samples are at a fixed distance in the taxonomy tree from the positive ones. To our knowledge, we are the first to demonstrate the ability to discriminate a single category from other semantically related ones using only its label.
title LLM meets Vision-Language Models for Zero-Shot One-Class Classification
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2404.00675