X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Swetha, Sirnam, Yang, Jinyu, Neiman, Tal, Rizve, Mamshad Nayeem, Tran, Son, Yao, Benjamin, Chilimbi, Trishul, Shah, Mubarak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916330049896448
author Swetha, Sirnam
Yang, Jinyu
Neiman, Tal
Rizve, Mamshad Nayeem
Tran, Son
Yao, Benjamin
Chilimbi, Trishul
Shah, Mubarak
author_facet Swetha, Sirnam
Yang, Jinyu
Neiman, Tal
Rizve, Mamshad Nayeem
Tran, Son
Yao, Benjamin
Chilimbi, Trishul
Shah, Mubarak
contents Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into Large Language Models (LLMs). The prevailing trend in this field involves the utilization of a vision encoder derived from vision-language contrastive learning (CL), showing expertise in capturing overall representations while facing difficulties in capturing detailed local patterns. In this work, we focus on enhancing the visual representations for MLLMs by combining high-frequency and detailed visual representations, obtained through masked image modeling (MIM), with semantically-enriched low-frequency representations captured by CL. To achieve this goal, we introduce X-Former which is a lightweight transformer module designed to exploit the complementary strengths of CL and MIM through an innovative interaction mechanism. Specifically, X-Former first bootstraps vision-language representation learning and multimodal-to-multimodal generative learning from two frozen vision encoders, i.e., CLIP-ViT (CL-based) and MAE-ViT (MIM-based). It further bootstraps vision-to-language generative learning from a frozen LLM to ensure visual features from X-Former can be interpreted by the LLM. To demonstrate the effectiveness of our approach, we assess its performance on tasks demanding detailed visual understanding. Extensive evaluations indicate that X-Former excels in visual reasoning tasks involving both structural and semantic categories in the GQA dataset. Assessment on fine-grained visual perception benchmark further confirms its superior capabilities in visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
Swetha, Sirnam
Yang, Jinyu
Neiman, Tal
Rizve, Mamshad Nayeem
Tran, Son
Yao, Benjamin
Chilimbi, Trishul
Shah, Mubarak
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into Large Language Models (LLMs). The prevailing trend in this field involves the utilization of a vision encoder derived from vision-language contrastive learning (CL), showing expertise in capturing overall representations while facing difficulties in capturing detailed local patterns. In this work, we focus on enhancing the visual representations for MLLMs by combining high-frequency and detailed visual representations, obtained through masked image modeling (MIM), with semantically-enriched low-frequency representations captured by CL. To achieve this goal, we introduce X-Former which is a lightweight transformer module designed to exploit the complementary strengths of CL and MIM through an innovative interaction mechanism. Specifically, X-Former first bootstraps vision-language representation learning and multimodal-to-multimodal generative learning from two frozen vision encoders, i.e., CLIP-ViT (CL-based) and MAE-ViT (MIM-based). It further bootstraps vision-to-language generative learning from a frozen LLM to ensure visual features from X-Former can be interpreted by the LLM. To demonstrate the effectiveness of our approach, we assess its performance on tasks demanding detailed visual understanding. Extensive evaluations indicate that X-Former excels in visual reasoning tasks involving both structural and semantic categories in the GQA dataset. Assessment on fine-grained visual perception benchmark further confirms its superior capabilities in visual understanding.
title X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
topic Computer Vision and Pattern Recognition
Machine Learning
Multimedia
url https://arxiv.org/abs/2407.13851