Why are Visually-Grounded Language Models Bad at Image Classification?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Yuhui, Unell, Alyssa, Wang, Xiaohan, Ghosh, Dhruba, Su, Yuchang, Schmidt, Ludwig, Yeung-Levy, Serena
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910681099403264
author Zhang, Yuhui
Unell, Alyssa
Wang, Xiaohan
Ghosh, Dhruba
Su, Yuchang
Schmidt, Ludwig
Yeung-Levy, Serena
author_facet Zhang, Yuhui
Unell, Alyssa
Wang, Xiaohan
Ghosh, Dhruba
Su, Yuchang
Schmidt, Ludwig
Yeung-Levy, Serena
contents Image classification is one of the most fundamental capabilities of machine vision intelligence. In this work, we revisit the image classification task using visually-grounded language models (VLMs) such as GPT-4V and LLaVA. We find that existing proprietary and public VLMs, despite often using CLIP as a vision encoder and having many more parameters, significantly underperform CLIP on standard image classification benchmarks like ImageNet. To understand the reason, we explore several hypotheses concerning the inference algorithms, training objectives, and data processing in VLMs. Our analysis reveals that the primary cause is data-related: critical information for image classification is encoded in the VLM's latent space but can only be effectively decoded with enough training data. Specifically, there is a strong correlation between the frequency of class exposure during VLM training and instruction-tuning and the VLM's performance in those classes; when trained with sufficient data, VLMs can match the accuracy of state-of-the-art classification models. Based on these findings, we enhance a VLM by integrating classification-focused datasets into its training, and demonstrate that the enhanced classification performance of the VLM transfers to its general capabilities, resulting in an improvement of 11.8% on the newly collected ImageWikiQA dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18415
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Why are Visually-Grounded Language Models Bad at Image Classification?
Zhang, Yuhui
Unell, Alyssa
Wang, Xiaohan
Ghosh, Dhruba
Su, Yuchang
Schmidt, Ludwig
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Image classification is one of the most fundamental capabilities of machine vision intelligence. In this work, we revisit the image classification task using visually-grounded language models (VLMs) such as GPT-4V and LLaVA. We find that existing proprietary and public VLMs, despite often using CLIP as a vision encoder and having many more parameters, significantly underperform CLIP on standard image classification benchmarks like ImageNet. To understand the reason, we explore several hypotheses concerning the inference algorithms, training objectives, and data processing in VLMs. Our analysis reveals that the primary cause is data-related: critical information for image classification is encoded in the VLM's latent space but can only be effectively decoded with enough training data. Specifically, there is a strong correlation between the frequency of class exposure during VLM training and instruction-tuning and the VLM's performance in those classes; when trained with sufficient data, VLMs can match the accuracy of state-of-the-art classification models. Based on these findings, we enhance a VLM by integrating classification-focused datasets into its training, and demonstrate that the enhanced classification performance of the VLM transfers to its general capabilities, resulting in an improvement of 11.8% on the newly collected ImageWikiQA dataset.
title Why are Visually-Grounded Language Models Bad at Image Classification?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.18415