Ranked from Within: Ranking Large Multimodal Models Without Labels
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918153379905536 |
|---|---|
| author | Tu, Weijie Deng, Weijian Campbell, Dylan Yao, Yu Zheng, Jiyang Gedeon, Tom Liu, Tongliang |
| author_facet | Tu, Weijie Deng, Weijian Campbell, Dylan Yao, Yu Zheng, Jiyang Gedeon, Tom Liu, Tongliang |
| contents | Can the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of giving the models an exam and marking them. We opt to avoid marking and the associated labor of determining the ground-truth answers. Instead, we explore other signals elicited and ascertain how well the models know their own limits, evaluating the effectiveness of these signals at unsupervised model ranking. We evaluate $47$ state-of-the-art LMMs (\eg, LLaVA) across $9$ visual question answering benchmarks, analyzing how well uncertainty-based metrics can predict relative model performance. Our findings show that uncertainty scores derived from softmax distributions provide a robust and consistent basis for ranking models across various tasks. This facilitates the ranking of LMMs on unlabeled data, providing a practical approach for selecting models for diverse target domains without requiring manual annotation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_06461 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Ranked from Within: Ranking Large Multimodal Models Without Labels Tu, Weijie Deng, Weijian Campbell, Dylan Yao, Yu Zheng, Jiyang Gedeon, Tom Liu, Tongliang Computer Vision and Pattern Recognition Can the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of giving the models an exam and marking them. We opt to avoid marking and the associated labor of determining the ground-truth answers. Instead, we explore other signals elicited and ascertain how well the models know their own limits, evaluating the effectiveness of these signals at unsupervised model ranking. We evaluate $47$ state-of-the-art LMMs (\eg, LLaVA) across $9$ visual question answering benchmarks, analyzing how well uncertainty-based metrics can predict relative model performance. Our findings show that uncertainty scores derived from softmax distributions provide a robust and consistent basis for ranking models across various tasks. This facilitates the ranking of LMMs on unlabeled data, providing a practical approach for selecting models for diverse target domains without requiring manual annotation. |
| title | Ranked from Within: Ranking Large Multimodal Models Without Labels |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.06461 |