Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Haowei, Shi, Yaya, Xu, Haiyang, Yuan, Chunfeng, Ye, Qinghao, Li, Chenliang, Yan, Ming, Zhang, Ji, Huang, Fei, Li, Bing, Hu, Weiming
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914698136387584
author Liu, Haowei
Shi, Yaya
Xu, Haiyang
Yuan, Chunfeng
Ye, Qinghao
Li, Chenliang
Yan, Ming
Zhang, Ji
Huang, Fei
Li, Bing
Hu, Weiming
author_facet Liu, Haowei
Shi, Yaya
Xu, Haiyang
Yuan, Chunfeng
Ye, Qinghao
Li, Chenliang
Yan, Ming
Zhang, Ji
Huang, Fei
Li, Bing
Hu, Weiming
contents In vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the reconstruction targets for MIM lack high-level semantics, and text is not sufficiently involved in masked modeling. These two drawbacks limit the effect of MIM in facilitating cross-modal semantic alignment. In this work, we propose a semantics-enhanced cross-modal MIM framework (SemMIM) for vision-language representation learning. Specifically, to provide more semantically meaningful supervision for MIM, we propose a local semantics enhancing approach, which harvest high-level semantics from global image features via self-supervised agreement learning and transfer them to local patch encodings by sharing the encoding space. Moreover, to achieve deep involvement of text during the entire MIM process, we propose a text-guided masking strategy and devise an efficient way of injecting textual information in both masked modeling and reconstruction target acquisition. Experimental results validate that our method improves the effectiveness of the MIM task in facilitating cross-modal semantic alignment. Compared to previous VLP models with similar model size and data scale, our SemMIM model achieves state-of-the-art or competitive performance on multiple downstream vision-language tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00249
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
Liu, Haowei
Shi, Yaya
Xu, Haiyang
Yuan, Chunfeng
Ye, Qinghao
Li, Chenliang
Yan, Ming
Zhang, Ji
Huang, Fei
Li, Bing
Hu, Weiming
Computer Vision and Pattern Recognition
In vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the reconstruction targets for MIM lack high-level semantics, and text is not sufficiently involved in masked modeling. These two drawbacks limit the effect of MIM in facilitating cross-modal semantic alignment. In this work, we propose a semantics-enhanced cross-modal MIM framework (SemMIM) for vision-language representation learning. Specifically, to provide more semantically meaningful supervision for MIM, we propose a local semantics enhancing approach, which harvest high-level semantics from global image features via self-supervised agreement learning and transfer them to local patch encodings by sharing the encoding space. Moreover, to achieve deep involvement of text during the entire MIM process, we propose a text-guided masking strategy and devise an efficient way of injecting textual information in both masked modeling and reconstruction target acquisition. Experimental results validate that our method improves the effectiveness of the MIM task in facilitating cross-modal semantic alignment. Compared to previous VLP models with similar model size and data scale, our SemMIM model achieves state-of-the-art or competitive performance on multiple downstream vision-language tasks.
title Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.00249