OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiayu, Jiao, Yang, Yu, Yue, Qian, Tianwen, Chen, Shaoxiang, Chen, Jingjing, Jiang, Yu-Gang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912392590393344
author Wang, Jiayu
Jiao, Yang
Yu, Yue
Qian, Tianwen
Chen, Shaoxiang
Chen, Jingjing
Jiang, Yu-Gang
author_facet Wang, Jiayu
Jiao, Yang
Yu, Yue
Qian, Tianwen
Chen, Shaoxiang
Chen, Jingjing
Jiang, Yu-Gang
contents Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack the necessary breadth and depth to fully evaluate the diverse capabilities of these models. To overcome this limitation, we introduce OmniGenBench, a novel and comprehensive benchmark meticulously designed to assess the instruction-following abilities of state-of-the-art LMMs across both perception-centric and cognition-centric dimensions. Our OmniGenBench includes 57 diverse sub-tasks grounded in real-world scenarios, systematically categorized according to the specific model capabilities they demand. For rigorous evaluation, we further employ a dual-mode protocol. This protocol utilizes off-the-shelf visual parsing tools for perception-centric tasks and a powerful LLM-based judger for cognition-centric tasks to assess the alignment between generated images and user instructions. Using OmniGenBench, we evaluate mainstream generative models, including prevalent models like GPT-4o, Gemini-2.0-Flash, and Seedream, and provide in-depth comparisons and analyses of their performance.Code and data are available at https://github.com/emilia113/OmniGenBench.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18775
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
Wang, Jiayu
Jiao, Yang
Yu, Yue
Qian, Tianwen
Chen, Shaoxiang
Chen, Jingjing
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack the necessary breadth and depth to fully evaluate the diverse capabilities of these models. To overcome this limitation, we introduce OmniGenBench, a novel and comprehensive benchmark meticulously designed to assess the instruction-following abilities of state-of-the-art LMMs across both perception-centric and cognition-centric dimensions. Our OmniGenBench includes 57 diverse sub-tasks grounded in real-world scenarios, systematically categorized according to the specific model capabilities they demand. For rigorous evaluation, we further employ a dual-mode protocol. This protocol utilizes off-the-shelf visual parsing tools for perception-centric tasks and a powerful LLM-based judger for cognition-centric tasks to assess the alignment between generated images and user instructions. Using OmniGenBench, we evaluate mainstream generative models, including prevalent models like GPT-4o, Gemini-2.0-Flash, and Seedream, and provide in-depth comparisons and analyses of their performance.Code and data are available at https://github.com/emilia113/OmniGenBench.
title OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.18775