MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Peng, Han, Siwei, Qiu, Shi, Zhou, Yiyang, Wang, Zhaoyang, Zheng, Wenhao, Chen, Zhaorun, Cui, Chenhang, Ding, Mingyu, Li, Linjie, Wang, Lijuan, Yao, Huaxiu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917972276150272
author Xia, Peng
Han, Siwei
Qiu, Shi
Zhou, Yiyang
Wang, Zhaoyang
Zheng, Wenhao
Chen, Zhaorun
Cui, Chenhang
Ding, Mingyu
Li, Linjie
Wang, Lijuan
Yao, Huaxiu
author_facet Xia, Peng
Han, Siwei
Qiu, Shi
Zhou, Yiyang
Wang, Zhaoyang
Zheng, Wenhao
Chen, Zhaorun
Cui, Chenhang
Ding, Mingyu
Li, Linjie
Wang, Lijuan
Yao, Huaxiu
contents Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation of this capability remains insufficient. Existing benchmarks suffer from limitations in data scale, scope, and evaluation depth, while current evaluation metrics are often costly or biased, lacking in reliability for practical applications. To address these challenges, we introduce MMIE, a large-scale knowledge-intensive benchmark for evaluating interleaved multimodal comprehension and generation in Large Vision-Language Models (LVLMs). MMIE comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. It supports both interleaved inputs and outputs, offering a mix of multiple-choice and open-ended question formats to evaluate diverse competencies. Moreover, we propose a reliable automated evaluation metric, leveraging a scoring model fine-tuned with human-annotated data and systematic evaluation criteria, aimed at reducing bias and improving evaluation accuracy. Extensive experiments demonstrate the effectiveness of our benchmark and metrics in providing a comprehensive evaluation of interleaved LVLMs. Specifically, we evaluate eight LVLMs, revealing that even the best models show significant room for improvement, with most achieving only moderate results. We believe MMIE will drive further advancements in the development of interleaved LVLMs. We publicly release our benchmark and code in https://mmie-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10139
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
Xia, Peng
Han, Siwei
Qiu, Shi
Zhou, Yiyang
Wang, Zhaoyang
Zheng, Wenhao
Chen, Zhaorun
Cui, Chenhang
Ding, Mingyu
Li, Linjie
Wang, Lijuan
Yao, Huaxiu
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation of this capability remains insufficient. Existing benchmarks suffer from limitations in data scale, scope, and evaluation depth, while current evaluation metrics are often costly or biased, lacking in reliability for practical applications. To address these challenges, we introduce MMIE, a large-scale knowledge-intensive benchmark for evaluating interleaved multimodal comprehension and generation in Large Vision-Language Models (LVLMs). MMIE comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. It supports both interleaved inputs and outputs, offering a mix of multiple-choice and open-ended question formats to evaluate diverse competencies. Moreover, we propose a reliable automated evaluation metric, leveraging a scoring model fine-tuned with human-annotated data and systematic evaluation criteria, aimed at reducing bias and improving evaluation accuracy. Extensive experiments demonstrate the effectiveness of our benchmark and metrics in providing a comprehensive evaluation of interleaved LVLMs. Specifically, we evaluate eight LVLMs, revealing that even the best models show significant room for improvement, with most achieving only moderate results. We believe MMIE will drive further advancements in the development of interleaved LVLMs. We publicly release our benchmark and code in https://mmie-bench.github.io/.
title MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.10139