Jina CLIP: Your CLIP Model Is Also Your Text Retriever
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917705235300352 |
|---|---|
| author | Koukounas, Andreas Mastrapas, Georgios Günther, Michael Wang, Bo Martens, Scott Mohr, Isabelle Sturua, Saba Akram, Mohammad Kalim Martínez, Joan Fontanals Ognawala, Saahil Guzman, Susana Werk, Maximilian Wang, Nan Xiao, Han |
| author_facet | Koukounas, Andreas Mastrapas, Georgios Günther, Michael Wang, Bo Martens, Scott Mohr, Isabelle Sturua, Saba Akram, Mohammad Kalim Martínez, Joan Fontanals Ognawala, Saahil Guzman, Susana Werk, Maximilian Wang, Nan Xiao, Han |
| contents | Contrastive Language-Image Pretraining (CLIP) is widely used to train models to align images and texts in a common embedding space by mapping them to fixed-sized vectors. These models are key to multimodal information retrieval and related tasks. However, CLIP models generally underperform in text-only tasks compared to specialized text models. This creates inefficiencies for information retrieval systems that keep separate embeddings and models for text-only and multimodal tasks. We propose a novel, multi-task contrastive training method to address this issue, which we use to train the jina-clip-v1 model to achieve the state-of-the-art performance on both text-image and text-text retrieval tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_20204 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Jina CLIP: Your CLIP Model Is Also Your Text Retriever Koukounas, Andreas Mastrapas, Georgios Günther, Michael Wang, Bo Martens, Scott Mohr, Isabelle Sturua, Saba Akram, Mohammad Kalim Martínez, Joan Fontanals Ognawala, Saahil Guzman, Susana Werk, Maximilian Wang, Nan Xiao, Han Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Information Retrieval 68T50 I.2.7 Contrastive Language-Image Pretraining (CLIP) is widely used to train models to align images and texts in a common embedding space by mapping them to fixed-sized vectors. These models are key to multimodal information retrieval and related tasks. However, CLIP models generally underperform in text-only tasks compared to specialized text models. This creates inefficiencies for information retrieval systems that keep separate embeddings and models for text-only and multimodal tasks. We propose a novel, multi-task contrastive training method to address this issue, which we use to train the jina-clip-v1 model to achieve the state-of-the-art performance on both text-image and text-text retrieval tasks. |
| title | Jina CLIP: Your CLIP Model Is Also Your Text Retriever |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Information Retrieval 68T50 I.2.7 |
| url | https://arxiv.org/abs/2405.20204 |