Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Yue, Zhang, Shuibai, Shao, Wenqi, Zhang, Kaipeng, Bin, Yi, Wang, Yu, Luo, Ping
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916758172991488
author Yang, Yue
Zhang, Shuibai
Shao, Wenqi
Zhang, Kaipeng
Bin, Yi
Wang, Yu
Luo, Ping
author_facet Yang, Yue
Zhang, Shuibai
Shao, Wenqi
Zhang, Kaipeng
Bin, Yi
Wang, Yu
Luo, Ping
contents Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across multimodal tasks such as visual perception and reasoning, leading to good performance on various multimodal evaluation benchmarks. However, these benchmarks keep a static nature and overlap with the pre-training data, resulting in fixed complexity constraints and data contamination issues. This raises the concern regarding the validity of the evaluation. To address these two challenges, we introduce a dynamic multimodal evaluation protocol called Vision-Language Bootstrapping (VLB). VLB provides a robust and comprehensive assessment for LVLMs with reduced data contamination and flexible complexity. To this end, VLB dynamically generates new visual question-answering samples through a multimodal bootstrapping module that modifies both images and language, while ensuring that newly generated samples remain consistent with the original ones by a judge module. By composing various bootstrapping strategies, VLB offers dynamic variants of existing benchmarks with diverse complexities, enabling the evaluation to co-evolve with the ever-evolving capabilities of LVLMs. Extensive experimental results across multiple benchmarks, including SEEDBench, MMBench, and MME, show that VLB significantly reduces data contamination and exposes performance limitations of LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08695
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
Yang, Yue
Zhang, Shuibai
Shao, Wenqi
Zhang, Kaipeng
Bin, Yi
Wang, Yu
Luo, Ping
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across multimodal tasks such as visual perception and reasoning, leading to good performance on various multimodal evaluation benchmarks. However, these benchmarks keep a static nature and overlap with the pre-training data, resulting in fixed complexity constraints and data contamination issues. This raises the concern regarding the validity of the evaluation. To address these two challenges, we introduce a dynamic multimodal evaluation protocol called Vision-Language Bootstrapping (VLB). VLB provides a robust and comprehensive assessment for LVLMs with reduced data contamination and flexible complexity. To this end, VLB dynamically generates new visual question-answering samples through a multimodal bootstrapping module that modifies both images and language, while ensuring that newly generated samples remain consistent with the original ones by a judge module. By composing various bootstrapping strategies, VLB offers dynamic variants of existing benchmarks with diverse complexities, enabling the evaluation to co-evolve with the ever-evolving capabilities of LVLMs. Extensive experimental results across multiple benchmarks, including SEEDBench, MMBench, and MME, show that VLB significantly reduces data contamination and exposes performance limitations of LVLMs.
title Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.08695