Exploiting LMM-based knowledge for image classification tasks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tzelepi, Maria, Mezaris, Vasileios
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914824794931200
author Tzelepi, Maria
Mezaris, Vasileios
author_facet Tzelepi, Maria
Mezaris, Vasileios
contents In this paper we address image classification tasks leveraging knowledge encoded in Large Multimodal Models (LMMs). More specifically, we use the MiniGPT-4 model to extract semantic descriptions for the images, in a multimodal prompting fashion. In the current literature, vision language models such as CLIP, among other approaches, are utilized as feature extractors, using only the image encoder, for solving image classification tasks. In this paper, we propose to additionally use the text encoder to obtain the text embeddings corresponding to the MiniGPT-4-generated semantic descriptions. Thus, we use both the image and text embeddings for solving the image classification task. The experimental evaluation on three datasets validates the improved classification performance achieved by exploiting LMM-based knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploiting LMM-based knowledge for image classification tasks
Tzelepi, Maria
Mezaris, Vasileios
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
In this paper we address image classification tasks leveraging knowledge encoded in Large Multimodal Models (LMMs). More specifically, we use the MiniGPT-4 model to extract semantic descriptions for the images, in a multimodal prompting fashion. In the current literature, vision language models such as CLIP, among other approaches, are utilized as feature extractors, using only the image encoder, for solving image classification tasks. In this paper, we propose to additionally use the text encoder to obtain the text embeddings corresponding to the MiniGPT-4-generated semantic descriptions. Thus, we use both the image and text embeddings for solving the image classification task. The experimental evaluation on three datasets validates the improved classification performance achieved by exploiting LMM-based knowledge.
title Exploiting LMM-based knowledge for image classification tasks
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2406.03071