MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Junjie, Liu, Zheng, Liu, Ze, Xiao, Shitao, Wang, Yueze, Zhao, Bo, Zhang, Chen Jason, Lian, Defu, Xiong, Yongping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915071072927744
author Zhou, Junjie
Liu, Zheng
Liu, Ze
Xiao, Shitao
Wang, Yueze
Zhao, Bo
Zhang, Chen Jason
Lian, Defu
Xiong, Yongping
author_facet Zhou, Junjie
Liu, Zheng
Liu, Ze
Xiao, Shitao
Wang, Yueze
Zhao, Bo
Zhang, Chen Jason
Lian, Defu
Xiong, Yongping
contents Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70$\times$ more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our produced dataset, well-trained models, and data synthesis pipeline will be made publicly available to facilitate the future development of this field.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14475
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Zhou, Junjie
Liu, Zheng
Liu, Ze
Xiao, Shitao
Wang, Yueze
Zhao, Bo
Zhang, Chen Jason
Lian, Defu
Xiong, Yongping
Computer Vision and Pattern Recognition
Computation and Language
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70$\times$ more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our produced dataset, well-trained models, and data synthesis pipeline will be made publicly available to facilitate the future development of this field.
title MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2412.14475