Better Language Models Exhibit Higher Visual Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruthardt, Jona, Burghouts, Gertjan J., Belongie, Serge, Asano, Yuki M.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908769145847808
author Ruthardt, Jona
Burghouts, Gertjan J.
Belongie, Serge
Asano, Yuki M.
author_facet Ruthardt, Jona
Burghouts, Gertjan J.
Belongie, Serge
Asano, Yuki M.
contents How well do text-only large language models (LLMs) align with the visual world? We present a systematic evaluation of this question by incorporating frozen representations of various language models into a discriminative vision-language framework and measuring zero-shot generalization to novel concepts. We find that decoder-based models exhibit stronger visual alignment than encoders, even when controlling for model and dataset size. Moreover, language modeling performance correlates with visual generalization, suggesting that advances in unimodal LLMs can simultaneously improve vision models. Leveraging these insights, we propose ShareLock, a lightweight method for fusing frozen vision and language backbones. ShareLock achieves robust performance across tasks while drastically reducing the need for paired data and compute. With just 563k image-caption pairs and under one GPU-hour of training, it reaches 51% accuracy on ImageNet. In cross-lingual settings, ShareLock dramatically outperforms CLIP, achieving 38.7% top-1 accuracy on Chinese image classification versus CLIP's 1.4%. Code is available.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07173
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Better Language Models Exhibit Higher Visual Alignment
Ruthardt, Jona
Burghouts, Gertjan J.
Belongie, Serge
Asano, Yuki M.
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
How well do text-only large language models (LLMs) align with the visual world? We present a systematic evaluation of this question by incorporating frozen representations of various language models into a discriminative vision-language framework and measuring zero-shot generalization to novel concepts. We find that decoder-based models exhibit stronger visual alignment than encoders, even when controlling for model and dataset size. Moreover, language modeling performance correlates with visual generalization, suggesting that advances in unimodal LLMs can simultaneously improve vision models. Leveraging these insights, we propose ShareLock, a lightweight method for fusing frozen vision and language backbones. ShareLock achieves robust performance across tasks while drastically reducing the need for paired data and compute. With just 563k image-caption pairs and under one GPU-hour of training, it reaches 51% accuracy on ImageNet. In cross-lingual settings, ShareLock dramatically outperforms CLIP, achieving 38.7% top-1 accuracy on Chinese image classification versus CLIP's 1.4%. Code is available.
title Better Language Models Exhibit Higher Visual Alignment
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.07173