LMM-Regularized CLIP Embeddings for Image Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tzelepi, Maria, Mezaris, Vasileios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917870088224768
author Tzelepi, Maria
Mezaris, Vasileios
author_facet Tzelepi, Maria
Mezaris, Vasileios
contents In this paper we deal with image classification tasks using the powerful CLIP vision-language model. Our goal is to advance the classification performance using the CLIP's image encoder, by proposing a novel Large Multimodal Model (LMM) based regularization method. The proposed method uses an LMM to extract semantic descriptions for the images of the dataset. Then, it uses the CLIP's text encoder, frozen, in order to obtain the corresponding text embeddings and compute the mean semantic class descriptions. Subsequently, we adapt the CLIP's image encoder by adding a classification head, and we train it along with the image encoder output, apart from the main classification objective, with an additional auxiliary objective. The additional objective forces the embeddings at the image encoder's output to become similar to their corresponding LMM-generated mean semantic class descriptions. In this way, it produces embeddings with enhanced discrimination ability, leading to improved classification performance. The effectiveness of the proposed regularization method is validated through extensive experiments on three image classification datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11663
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LMM-Regularized CLIP Embeddings for Image Classification
Tzelepi, Maria
Mezaris, Vasileios
Computer Vision and Pattern Recognition
Multimedia
In this paper we deal with image classification tasks using the powerful CLIP vision-language model. Our goal is to advance the classification performance using the CLIP's image encoder, by proposing a novel Large Multimodal Model (LMM) based regularization method. The proposed method uses an LMM to extract semantic descriptions for the images of the dataset. Then, it uses the CLIP's text encoder, frozen, in order to obtain the corresponding text embeddings and compute the mean semantic class descriptions. Subsequently, we adapt the CLIP's image encoder by adding a classification head, and we train it along with the image encoder output, apart from the main classification objective, with an additional auxiliary objective. The additional objective forces the embeddings at the image encoder's output to become similar to their corresponding LMM-generated mean semantic class descriptions. In this way, it produces embeddings with enhanced discrimination ability, leading to improved classification performance. The effectiveness of the proposed regularization method is validated through extensive experiments on three image classification datasets.
title LMM-Regularized CLIP Embeddings for Image Classification
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2412.11663