Do Vision and Language Encoders Represent the World Similarly?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Maniparambil, Mayug, Akshulakov, Raiymbek, Djilali, Yasser Abdelaziz Dahou, Narayan, Sanath, Seddik, Mohamed El Amine, Mangalam, Karttikeya, O'Connor, Noel E.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911810299363328
author Maniparambil, Mayug
Akshulakov, Raiymbek
Djilali, Yasser Abdelaziz Dahou
Narayan, Sanath
Seddik, Mohamed El Amine
Mangalam, Karttikeya
O'Connor, Noel E.
author_facet Maniparambil, Mayug
Akshulakov, Raiymbek
Djilali, Yasser Abdelaziz Dahou
Narayan, Sanath
Seddik, Mohamed El Amine
Mangalam, Karttikeya
O'Connor, Noel E.
contents Aligned text-image encoders such as CLIP have become the de facto model for vision-language tasks. Furthermore, modality-specific encoders achieve impressive performances in their respective domains. This raises a central question: does an alignment exist between uni-modal vision and language encoders since they fundamentally represent the same physical world? Analyzing the latent spaces structure of vision and language models on image-caption benchmarks using the Centered Kernel Alignment (CKA), we find that the representation spaces of unaligned and aligned encoders are semantically similar. In the absence of statistical similarity in aligned encoders like CLIP, we show that a possible matching of unaligned encoders exists without any training. We frame this as a seeded graph-matching problem exploiting the semantic similarity between graphs and propose two methods - a Fast Quadratic Assignment Problem optimization, and a novel localized CKA metric-based matching/retrieval. We demonstrate the effectiveness of this on several downstream tasks including cross-lingual, cross-domain caption matching and image classification. Code available at github.com/mayug/0-shot-llm-vision.
format Preprint
id arxiv_https___arxiv_org_abs_2401_05224
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do Vision and Language Encoders Represent the World Similarly?
Maniparambil, Mayug
Akshulakov, Raiymbek
Djilali, Yasser Abdelaziz Dahou
Narayan, Sanath
Seddik, Mohamed El Amine
Mangalam, Karttikeya
O'Connor, Noel E.
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Aligned text-image encoders such as CLIP have become the de facto model for vision-language tasks. Furthermore, modality-specific encoders achieve impressive performances in their respective domains. This raises a central question: does an alignment exist between uni-modal vision and language encoders since they fundamentally represent the same physical world? Analyzing the latent spaces structure of vision and language models on image-caption benchmarks using the Centered Kernel Alignment (CKA), we find that the representation spaces of unaligned and aligned encoders are semantically similar. In the absence of statistical similarity in aligned encoders like CLIP, we show that a possible matching of unaligned encoders exists without any training. We frame this as a seeded graph-matching problem exploiting the semantic similarity between graphs and propose two methods - a Fast Quadratic Assignment Problem optimization, and a novel localized CKA metric-based matching/retrieval. We demonstrate the effectiveness of this on several downstream tasks including cross-lingual, cross-domain caption matching and image classification. Code available at github.com/mayug/0-shot-llm-vision.
title Do Vision and Language Encoders Represent the World Similarly?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.05224