Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.16006 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913328127803392 |
|---|---|
| author | Ying, Kaining Meng, Fanqing Wang, Jin Li, Zhiqian Lin, Han Yang, Yue Zhang, Hao Zhang, Wenbo Lin, Yuqi Liu, Shuo Lei, Jiayi Lu, Quanfeng Chen, Runjian Xu, Peng Zhang, Renrui Zhang, Haozhe Gao, Peng Wang, Yali Qiao, Yu Luo, Ping Zhang, Kaipeng Shao, Wenqi |
| author_facet | Ying, Kaining Meng, Fanqing Wang, Jin Li, Zhiqian Lin, Han Yang, Yue Zhang, Hao Zhang, Wenbo Lin, Yuqi Liu, Shuo Lei, Jiayi Lu, Quanfeng Chen, Runjian Xu, Peng Zhang, Renrui Zhang, Haozhe Gao, Peng Wang, Yali Qiao, Yu Luo, Ping Zhang, Kaipeng Shao, Wenqi |
| contents | Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, reasoning, and planning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $30$ LVLMs such as the proprietary GPT-4V, GeminiProVision, and open-sourced InternVL-Chat, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_16006 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI Ying, Kaining Meng, Fanqing Wang, Jin Li, Zhiqian Lin, Han Yang, Yue Zhang, Hao Zhang, Wenbo Lin, Yuqi Liu, Shuo Lei, Jiayi Lu, Quanfeng Chen, Runjian Xu, Peng Zhang, Renrui Zhang, Haozhe Gao, Peng Wang, Yali Qiao, Yu Luo, Ping Zhang, Kaipeng Shao, Wenqi Computer Vision and Pattern Recognition Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, reasoning, and planning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $30$ LVLMs such as the proprietary GPT-4V, GeminiProVision, and open-sourced InternVL-Chat, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence. |
| title | MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2404.16006 |