MLLMs-Augmented Visual-Language Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yanqing, Wang, Kai, Shao, Wenqi, Luo, Ping, Qiao, Yu, Shou, Mike Zheng, Zhang, Kaipeng, You, Yang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911795566870528
author Liu, Yanqing
Wang, Kai
Shao, Wenqi
Luo, Ping
Qiao, Yu
Shou, Mike Zheng
Zhang, Kaipeng
You, Yang
author_facet Liu, Yanqing
Wang, Kai
Shao, Wenqi
Luo, Ping
Qiao, Yu
Shou, Mike Zheng
Zhang, Kaipeng
You, Yang
contents Visual-language pre-training has achieved remarkable success in many multi-modal tasks, largely attributed to the availability of large-scale image-text datasets. In this work, we demonstrate that Multi-modal Large Language Models (MLLMs) can enhance visual-language representation learning by establishing richer image-text associations for image-text datasets. Our approach is simple, utilizing MLLMs to extend multiple diverse captions for each image. To prevent the bias introduced by MLLMs' hallucinations and monotonous language styles, we propose "text shearing" to maintain the quality and availability of extended captions. In image-text retrieval, without introducing additional training cost, our method consistently obtains 5.6 ~ 35.0 and 16.8 ~ 46.1 improvement on Recall@1 under the fine-tuning and zero-shot settings, respectively. Notably, we obtain zero-shot results that are comparable to fine-tuning on target datasets, which encourages more exploration of the versatile use of MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18765
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MLLMs-Augmented Visual-Language Representation Learning
Liu, Yanqing
Wang, Kai
Shao, Wenqi
Luo, Ping
Qiao, Yu
Shou, Mike Zheng
Zhang, Kaipeng
You, Yang
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Visual-language pre-training has achieved remarkable success in many multi-modal tasks, largely attributed to the availability of large-scale image-text datasets. In this work, we demonstrate that Multi-modal Large Language Models (MLLMs) can enhance visual-language representation learning by establishing richer image-text associations for image-text datasets. Our approach is simple, utilizing MLLMs to extend multiple diverse captions for each image. To prevent the bias introduced by MLLMs' hallucinations and monotonous language styles, we propose "text shearing" to maintain the quality and availability of extended captions. In image-text retrieval, without introducing additional training cost, our method consistently obtains 5.6 ~ 35.0 and 16.8 ~ 46.1 improvement on Recall@1 under the fine-tuning and zero-shot settings, respectively. Notably, we obtain zero-shot results that are comparable to fine-tuning on target datasets, which encourages more exploration of the versatile use of MLLMs.
title MLLMs-Augmented Visual-Language Representation Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2311.18765