MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Jihan, Hu, Yushi, Yi, Yujie, Han, Bin, Feng, Shangbin, Yang, Guang, Wen, Bingbing, Krishna, Ranjay, Wang, Lucy Lu, Tsvetkov, Yulia, Smith, Noah A., Zhu, Banghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916754743099392
author Yao, Jihan
Hu, Yushi
Yi, Yujie
Han, Bin
Feng, Shangbin
Yang, Guang
Wen, Bingbing
Krishna, Ranjay
Wang, Lucy Lu
Tsvetkov, Yulia
Smith, Noah A.
Zhu, Banghua
author_facet Yao, Jihan
Hu, Yushi
Yi, Yujie
Han, Bin
Feng, Shangbin
Yang, Guang
Wen, Bingbing
Krishna, Ranjay
Wang, Lucy Lu
Tsvetkov, Yulia
Smith, Noah A.
Zhu, Banghua
contents Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we present MMMG, a comprehensive and human-aligned benchmark for multimodal generation across 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio), with a focus on tasks that present significant challenges for generation models, while still enabling reliable automatic evaluation through a combination of models and programs. MMMG encompasses 49 tasks (including 29 newly developed ones), each with a carefully designed evaluation pipeline, and 937 instructions to systematically assess reasoning, controllability, and other key capabilities of multimodal generation models. Extensive validation demonstrates that MMMG is highly aligned with human evaluation, achieving an average agreement of 94.3%. Benchmarking results on 24 multimodal generation models reveal that even though the state-of-the-art model, GPT Image, achieves 78.3% accuracy for image generation, it falls short on multimodal reasoning and interleaved generation. Furthermore, results suggest considerable headroom for improvement in audio generation, highlighting an important direction for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
Yao, Jihan
Hu, Yushi
Yi, Yujie
Han, Bin
Feng, Shangbin
Yang, Guang
Wen, Bingbing
Krishna, Ranjay
Wang, Lucy Lu
Tsvetkov, Yulia
Smith, Noah A.
Zhu, Banghua
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we present MMMG, a comprehensive and human-aligned benchmark for multimodal generation across 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio), with a focus on tasks that present significant challenges for generation models, while still enabling reliable automatic evaluation through a combination of models and programs. MMMG encompasses 49 tasks (including 29 newly developed ones), each with a carefully designed evaluation pipeline, and 937 instructions to systematically assess reasoning, controllability, and other key capabilities of multimodal generation models. Extensive validation demonstrates that MMMG is highly aligned with human evaluation, achieving an average agreement of 94.3%. Benchmarking results on 24 multimodal generation models reveal that even though the state-of-the-art model, GPT Image, achieves 78.3% accuracy for image generation, it falls short on multimodal reasoning and interleaved generation. Furthermore, results suggest considerable headroom for improvement in audio generation, highlighting an important direction for future research.
title MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17613