Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koukounas, Andreas, Mastrapas, Georgios, Günther, Michael, Wang, Bo, Martens, Scott, Mohr, Isabelle, Sturua, Saba, Akram, Mohammad Kalim, Martínez, Joan Fontanals, Ognawala, Saahil, Guzman, Susana, Werk, Maximilian, Wang, Nan, Xiao, Han
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917705235300352
author Koukounas, Andreas
Mastrapas, Georgios
Günther, Michael
Wang, Bo
Martens, Scott
Mohr, Isabelle
Sturua, Saba
Akram, Mohammad Kalim
Martínez, Joan Fontanals
Ognawala, Saahil
Guzman, Susana
Werk, Maximilian
Wang, Nan
Xiao, Han
author_facet Koukounas, Andreas
Mastrapas, Georgios
Günther, Michael
Wang, Bo
Martens, Scott
Mohr, Isabelle
Sturua, Saba
Akram, Mohammad Kalim
Martínez, Joan Fontanals
Ognawala, Saahil
Guzman, Susana
Werk, Maximilian
Wang, Nan
Xiao, Han
contents Contrastive Language-Image Pretraining (CLIP) is widely used to train models to align images and texts in a common embedding space by mapping them to fixed-sized vectors. These models are key to multimodal information retrieval and related tasks. However, CLIP models generally underperform in text-only tasks compared to specialized text models. This creates inefficiencies for information retrieval systems that keep separate embeddings and models for text-only and multimodal tasks. We propose a novel, multi-task contrastive training method to address this issue, which we use to train the jina-clip-v1 model to achieve the state-of-the-art performance on both text-image and text-text retrieval tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20204
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Jina CLIP: Your CLIP Model Is Also Your Text Retriever
Koukounas, Andreas
Mastrapas, Georgios
Günther, Michael
Wang, Bo
Martens, Scott
Mohr, Isabelle
Sturua, Saba
Akram, Mohammad Kalim
Martínez, Joan Fontanals
Ognawala, Saahil
Guzman, Susana
Werk, Maximilian
Wang, Nan
Xiao, Han
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
68T50
I.2.7
Contrastive Language-Image Pretraining (CLIP) is widely used to train models to align images and texts in a common embedding space by mapping them to fixed-sized vectors. These models are key to multimodal information retrieval and related tasks. However, CLIP models generally underperform in text-only tasks compared to specialized text models. This creates inefficiencies for information retrieval systems that keep separate embeddings and models for text-only and multimodal tasks. We propose a novel, multi-task contrastive training method to address this issue, which we use to train the jina-clip-v1 model to achieve the state-of-the-art performance on both text-image and text-text retrieval tasks.
title Jina CLIP: Your CLIP Model Is Also Your Text Retriever
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
68T50
I.2.7
url https://arxiv.org/abs/2405.20204