Saved in:
Bibliographic Details
Main Authors: Tiong, Anthony Meng Huat, Zhao, Junqi, Li, Boyang, Li, Junnan, Hoi, Steven C. H., Xiong, Caiming
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.02415
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914739375833088
author Tiong, Anthony Meng Huat
Zhao, Junqi
Li, Boyang
Li, Junnan
Hoi, Steven C. H.
Xiong, Caiming
author_facet Tiong, Anthony Meng Huat
Zhao, Junqi
Li, Boyang
Li, Junnan
Hoi, Steven C. H.
Xiong, Caiming
contents Vision-language (VL) models, pretrained on colossal image-text datasets, have attained broad VL competence that is difficult to evaluate. A common belief is that a small number of VL skills underlie the variety of VL tests. In this paper, we perform a large-scale transfer learning experiment aimed at discovering latent VL skills from data. We reveal interesting characteristics that have important implications for test suite design. First, generation tasks suffer from a length bias, suggesting benchmarks should balance tasks with varying output lengths. Second, we demonstrate that factor analysis successfully identifies reasonable yet surprising VL skill factors, suggesting benchmarks could leverage similar analyses for task selection. Finally, we present a new dataset, OLIVE (https://github.com/jq-zh/olive-dataset), which simulates user instructions in the wild and presents challenges dissimilar to all datasets we tested. Our findings contribute to the design of balanced and broad-coverage vision-language evaluation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02415
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
Tiong, Anthony Meng Huat
Zhao, Junqi
Li, Boyang
Li, Junnan
Hoi, Steven C. H.
Xiong, Caiming
Computer Vision and Pattern Recognition
Vision-language (VL) models, pretrained on colossal image-text datasets, have attained broad VL competence that is difficult to evaluate. A common belief is that a small number of VL skills underlie the variety of VL tests. In this paper, we perform a large-scale transfer learning experiment aimed at discovering latent VL skills from data. We reveal interesting characteristics that have important implications for test suite design. First, generation tasks suffer from a length bias, suggesting benchmarks should balance tasks with varying output lengths. Second, we demonstrate that factor analysis successfully identifies reasonable yet surprising VL skill factors, suggesting benchmarks could leverage similar analyses for task selection. Finally, we present a new dataset, OLIVE (https://github.com/jq-zh/olive-dataset), which simulates user instructions in the wild and presents challenges dissimilar to all datasets we tested. Our findings contribute to the design of balanced and broad-coverage vision-language evaluation methods.
title What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.02415