COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Sanghwan, Xiao, Rui, Georgescu, Mariana-Iuliana, Alaniz, Stephan, Akata, Zeynep
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916662581657600
author Kim, Sanghwan
Xiao, Rui
Georgescu, Mariana-Iuliana
Alaniz, Stephan
Akata, Zeynep
author_facet Kim, Sanghwan
Xiao, Rui
Georgescu, Mariana-Iuliana
Alaniz, Stephan
Akata, Zeynep
contents Vision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly on foreground objects, neglecting other crucial information in the image, which limits their effectiveness in downstream tasks. To address these challenges, we propose COSMOS: CrOSs-MOdality Self-distillation for vision-language pre-training that integrates a novel text-cropping strategy and cross-attention module into a self-supervised learning framework. We create global and local views of images and texts (i.e., multi-modal augmentations), which are essential for self-distillation in VLMs. We further introduce a cross-attention module, enabling COSMOS to learn comprehensive cross-modal representations optimized via a cross-modality self-distillation loss. COSMOS consistently outperforms previous strong baselines on various zero-shot downstream tasks, including retrieval, classification, and semantic segmentation. Additionally, it surpasses CLIP-based models trained on larger datasets in visual perception and contextual understanding tasks. Code is available at https://github.com/ExplainableML/cosmos.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
Kim, Sanghwan
Xiao, Rui
Georgescu, Mariana-Iuliana
Alaniz, Stephan
Akata, Zeynep
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Vision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly on foreground objects, neglecting other crucial information in the image, which limits their effectiveness in downstream tasks. To address these challenges, we propose COSMOS: CrOSs-MOdality Self-distillation for vision-language pre-training that integrates a novel text-cropping strategy and cross-attention module into a self-supervised learning framework. We create global and local views of images and texts (i.e., multi-modal augmentations), which are essential for self-distillation in VLMs. We further introduce a cross-attention module, enabling COSMOS to learn comprehensive cross-modal representations optimized via a cross-modality self-distillation loss. COSMOS consistently outperforms previous strong baselines on various zero-shot downstream tasks, including retrieval, classification, and semantic segmentation. Additionally, it surpasses CLIP-based models trained on larger datasets in visual perception and contextual understanding tasks. Code is available at https://github.com/ExplainableML/cosmos.
title COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.01814