Saved in:
Bibliographic Details
Main Authors: Ying, Kaining, Meng, Fanqing, Wang, Jin, Li, Zhiqian, Lin, Han, Yang, Yue, Zhang, Hao, Zhang, Wenbo, Lin, Yuqi, Liu, Shuo, Lei, Jiayi, Lu, Quanfeng, Chen, Runjian, Xu, Peng, Zhang, Renrui, Zhang, Haozhe, Gao, Peng, Wang, Yali, Qiao, Yu, Luo, Ping, Zhang, Kaipeng, Shao, Wenqi
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.16006
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913328127803392
author Ying, Kaining
Meng, Fanqing
Wang, Jin
Li, Zhiqian
Lin, Han
Yang, Yue
Zhang, Hao
Zhang, Wenbo
Lin, Yuqi
Liu, Shuo
Lei, Jiayi
Lu, Quanfeng
Chen, Runjian
Xu, Peng
Zhang, Renrui
Zhang, Haozhe
Gao, Peng
Wang, Yali
Qiao, Yu
Luo, Ping
Zhang, Kaipeng
Shao, Wenqi
author_facet Ying, Kaining
Meng, Fanqing
Wang, Jin
Li, Zhiqian
Lin, Han
Yang, Yue
Zhang, Hao
Zhang, Wenbo
Lin, Yuqi
Liu, Shuo
Lei, Jiayi
Lu, Quanfeng
Chen, Runjian
Xu, Peng
Zhang, Renrui
Zhang, Haozhe
Gao, Peng
Wang, Yali
Qiao, Yu
Luo, Ping
Zhang, Kaipeng
Shao, Wenqi
contents Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, reasoning, and planning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $30$ LVLMs such as the proprietary GPT-4V, GeminiProVision, and open-sourced InternVL-Chat, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16006
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Ying, Kaining
Meng, Fanqing
Wang, Jin
Li, Zhiqian
Lin, Han
Yang, Yue
Zhang, Hao
Zhang, Wenbo
Lin, Yuqi
Liu, Shuo
Lei, Jiayi
Lu, Quanfeng
Chen, Runjian
Xu, Peng
Zhang, Renrui
Zhang, Haozhe
Gao, Peng
Wang, Yali
Qiao, Yu
Luo, Ping
Zhang, Kaipeng
Shao, Wenqi
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, reasoning, and planning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $30$ LVLMs such as the proprietary GPT-4V, GeminiProVision, and open-sourced InternVL-Chat, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence.
title MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.16006