UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yi, Wang, Haonan, Zhang, Qixiang, Xiao, Boyu, Hu, Chenchang, Wang, Hualiang, Li, Xiaomeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918021661982720
author Li, Yi
Wang, Haonan
Zhang, Qixiang
Xiao, Boyu
Hu, Chenchang
Wang, Hualiang
Li, Xiaomeng
author_facet Li, Yi
Wang, Haonan
Zhang, Qixiang
Xiao, Boyu
Hu, Chenchang
Wang, Hualiang
Li, Xiaomeng
contents The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, there is a lack of a unified evaluation framework for these models, which would enable an elegant, simplified, and overall evaluation. Current models conduct evaluations on multiple task-specific benchmarks, but there are significant limitations, such as the lack of overall results, errors from extra evaluation models, reliance on extensive labeled images, benchmarks that lack diversity, and metrics with limited capacity for instruction-following evaluation. To tackle these challenges, we introduce UniEval, the first evaluation framework designed for unified multimodal models without extra models, images, or annotations. This facilitates a simplified and unified evaluation process. The UniEval framework contains a holistic benchmark, UniBench (supports both unified and visual generation models), along with the corresponding UniScore metric. UniBench includes 81 fine-grained tags contributing to high diversity. Experimental results indicate that UniBench is more challenging than existing benchmarks, and UniScore aligns closely with human evaluations, surpassing current metrics. Moreover, we extensively evaluated SoTA unified and visual generation models, uncovering new insights into Univeral's unique values.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
Li, Yi
Wang, Haonan
Zhang, Qixiang
Xiao, Boyu
Hu, Chenchang
Wang, Hualiang
Li, Xiaomeng
Computer Vision and Pattern Recognition
Artificial Intelligence
The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, there is a lack of a unified evaluation framework for these models, which would enable an elegant, simplified, and overall evaluation. Current models conduct evaluations on multiple task-specific benchmarks, but there are significant limitations, such as the lack of overall results, errors from extra evaluation models, reliance on extensive labeled images, benchmarks that lack diversity, and metrics with limited capacity for instruction-following evaluation. To tackle these challenges, we introduce UniEval, the first evaluation framework designed for unified multimodal models without extra models, images, or annotations. This facilitates a simplified and unified evaluation process. The UniEval framework contains a holistic benchmark, UniBench (supports both unified and visual generation models), along with the corresponding UniScore metric. UniBench includes 81 fine-grained tags contributing to high diversity. Experimental results indicate that UniBench is more challenging than existing benchmarks, and UniScore aligns closely with human evaluations, surpassing current metrics. Moreover, we extensively evaluated SoTA unified and visual generation models, uncovering new insights into Univeral's unique values.
title UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.10483