Enhancing Image Retrieval : A Comprehensive Study on Photo Search using the CLIP Mode

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lahajal, Naresh Kumar, S, Harini
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916104422555648
author Lahajal, Naresh Kumar
S, Harini
author_facet Lahajal, Naresh Kumar
S, Harini
contents Photo search, the task of retrieving images based on textual queries, has witnessed significant advancements with the introduction of CLIP (Contrastive Language-Image Pretraining) model. CLIP leverages a vision-language pre training approach, wherein it learns a shared representation space for images and text, enabling cross-modal understanding. This model demonstrates the capability to understand the semantic relationships between diverse image and text pairs, allowing for efficient and accurate retrieval of images based on natural language queries. By training on a large-scale dataset containing images and their associated textual descriptions, CLIP achieves remarkable generalization, providing a powerful tool for tasks such as zero-shot learning and few-shot classification. This abstract summarizes the foundational principles of CLIP and highlights its potential impact on advancing the field of photo search, fostering a seamless integration of natural language understanding and computer vision for improved information retrieval in multimedia applications
format Preprint
id arxiv_https___arxiv_org_abs_2401_13613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Image Retrieval : A Comprehensive Study on Photo Search using the CLIP Mode
Lahajal, Naresh Kumar
S, Harini
Computer Vision and Pattern Recognition
Artificial Intelligence
Photo search, the task of retrieving images based on textual queries, has witnessed significant advancements with the introduction of CLIP (Contrastive Language-Image Pretraining) model. CLIP leverages a vision-language pre training approach, wherein it learns a shared representation space for images and text, enabling cross-modal understanding. This model demonstrates the capability to understand the semantic relationships between diverse image and text pairs, allowing for efficient and accurate retrieval of images based on natural language queries. By training on a large-scale dataset containing images and their associated textual descriptions, CLIP achieves remarkable generalization, providing a powerful tool for tasks such as zero-shot learning and few-shot classification. This abstract summarizes the foundational principles of CLIP and highlights its potential impact on advancing the field of photo search, fostering a seamless integration of natural language understanding and computer vision for improved information retrieval in multimedia applications
title Enhancing Image Retrieval : A Comprehensive Study on Photo Search using the CLIP Mode
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2401.13613