Ranked from Within: Ranking Large Multimodal Models Without Labels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Weijie, Deng, Weijian, Campbell, Dylan, Yao, Yu, Zheng, Jiyang, Gedeon, Tom, Liu, Tongliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918153379905536
author Tu, Weijie
Deng, Weijian
Campbell, Dylan
Yao, Yu
Zheng, Jiyang
Gedeon, Tom
Liu, Tongliang
author_facet Tu, Weijie
Deng, Weijian
Campbell, Dylan
Yao, Yu
Zheng, Jiyang
Gedeon, Tom
Liu, Tongliang
contents Can the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of giving the models an exam and marking them. We opt to avoid marking and the associated labor of determining the ground-truth answers. Instead, we explore other signals elicited and ascertain how well the models know their own limits, evaluating the effectiveness of these signals at unsupervised model ranking. We evaluate $47$ state-of-the-art LMMs (\eg, LLaVA) across $9$ visual question answering benchmarks, analyzing how well uncertainty-based metrics can predict relative model performance. Our findings show that uncertainty scores derived from softmax distributions provide a robust and consistent basis for ranking models across various tasks. This facilitates the ranking of LMMs on unlabeled data, providing a practical approach for selecting models for diverse target domains without requiring manual annotation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06461
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ranked from Within: Ranking Large Multimodal Models Without Labels
Tu, Weijie
Deng, Weijian
Campbell, Dylan
Yao, Yu
Zheng, Jiyang
Gedeon, Tom
Liu, Tongliang
Computer Vision and Pattern Recognition
Can the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of giving the models an exam and marking them. We opt to avoid marking and the associated labor of determining the ground-truth answers. Instead, we explore other signals elicited and ascertain how well the models know their own limits, evaluating the effectiveness of these signals at unsupervised model ranking. We evaluate $47$ state-of-the-art LMMs (\eg, LLaVA) across $9$ visual question answering benchmarks, analyzing how well uncertainty-based metrics can predict relative model performance. Our findings show that uncertainty scores derived from softmax distributions provide a robust and consistent basis for ranking models across various tasks. This facilitates the ranking of LMMs on unlabeled data, providing a practical approach for selecting models for diverse target domains without requiring manual annotation.
title Ranked from Within: Ranking Large Multimodal Models Without Labels
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.06461