LIME: Less Is More for MLLM Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, King, Zang, Qianbo, Jia, Shian, Wu, Siwei, Fang, Feiteng, Li, Yizhi, Gavin, Shawn, Zheng, Tuney, Guo, Jiawei, Li, Bo, Wu, Haoning, Qu, Xingwei, Yang, Jian, Liu, Zachary, Yue, Xiang, Liu, J. H., Lin, Chenghua, Yang, Min, Ni, Shiwen, Huang, Wenhao, Zhang, Ge
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912071026737152
author Zhu, King
Zang, Qianbo
Jia, Shian
Wu, Siwei
Fang, Feiteng
Li, Yizhi
Gavin, Shawn
Zheng, Tuney
Guo, Jiawei
Li, Bo
Wu, Haoning
Qu, Xingwei
Yang, Jian
Liu, Zachary
Yue, Xiang
Liu, J. H.
Lin, Chenghua
Yang, Min
Ni, Shiwen
Huang, Wenhao
Zhang, Ge
author_facet Zhu, King
Zang, Qianbo
Jia, Shian
Wu, Siwei
Fang, Feiteng
Li, Yizhi
Gavin, Shawn
Zheng, Tuney
Guo, Jiawei
Li, Bo
Wu, Haoning
Qu, Xingwei
Yang, Jian
Liu, Zachary
Yue, Xiang
Liu, J. H.
Lin, Chenghua
Yang, Min
Ni, Shiwen
Huang, Wenhao
Zhang, Ge
contents Multimodal Large Language Models (MLLMs) are evaluated on various benchmarks, such as image captioning, visual question answering, and reasoning. However, many of these benchmarks include overly simple or uninformative samples, complicating the effective distinction of different MLLMs' performance. Furthermore, evaluating models across numerous benchmarks incurs a significant computational burden. To address these issues, we propose LIME (Less Is More for MLLM Evaluation), a refined and efficient benchmark curated through a semi-automated pipeline. This pipeline filters out uninformative samples and eliminates answer leakage by focusing on tasks that necessitate image-based understanding. Our experiments indicate that LIME reduces the number of samples by 76% and evaluation time by 77%, while also providing a more effective means of distinguishing the capabilities of different models. Notably, we find that traditional automatic metrics, such as CIDEr, are inadequate for assessing MLLMs' captioning performance; excluding the caption task score yields a more accurate reflection of overall model performance. All code and data are available at https://github.com/kangreen0210/LIME.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LIME: Less Is More for MLLM Evaluation
Zhu, King
Zang, Qianbo
Jia, Shian
Wu, Siwei
Fang, Feiteng
Li, Yizhi
Gavin, Shawn
Zheng, Tuney
Guo, Jiawei
Li, Bo
Wu, Haoning
Qu, Xingwei
Yang, Jian
Liu, Zachary
Yue, Xiang
Liu, J. H.
Lin, Chenghua
Yang, Min
Ni, Shiwen
Huang, Wenhao
Zhang, Ge
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal Large Language Models (MLLMs) are evaluated on various benchmarks, such as image captioning, visual question answering, and reasoning. However, many of these benchmarks include overly simple or uninformative samples, complicating the effective distinction of different MLLMs' performance. Furthermore, evaluating models across numerous benchmarks incurs a significant computational burden. To address these issues, we propose LIME (Less Is More for MLLM Evaluation), a refined and efficient benchmark curated through a semi-automated pipeline. This pipeline filters out uninformative samples and eliminates answer leakage by focusing on tasks that necessitate image-based understanding. Our experiments indicate that LIME reduces the number of samples by 76% and evaluation time by 77%, while also providing a more effective means of distinguishing the capabilities of different models. Notably, we find that traditional automatic metrics, such as CIDEr, are inadequate for assessing MLLMs' captioning performance; excluding the caption task score yields a more accurate reflection of overall model performance. All code and data are available at https://github.com/kangreen0210/LIME.
title LIME: Less Is More for MLLM Evaluation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2409.06851