GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Ruihang, Qu, Leigang, Zhang, Jingxu, Gui, Dongnan, Xu, Mengde, Zhang, Xiaosong, Hu, Han, Wang, Wenjie, Wang, Jiaqi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914308543217664
author Li, Ruihang
Qu, Leigang
Zhang, Jingxu
Gui, Dongnan
Xu, Mengde
Zhang, Xiaosong
Hu, Han
Wang, Wenjie
Wang, Jiaqi
author_facet Li, Ruihang
Qu, Leigang
Zhang, Jingxu
Gui, Dongnan
Xu, Mengde
Zhang, Xiaosong
Hu, Han
Wang, Wenjie
Wang, Jiaqi
contents The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this work, we systematically investigate the reliability of the prevailing absolute pointwise scoring standard, across a wide spectrum of visual generation tasks. Our analysis reveals that this paradigm is limited due to stochastic inconsistency and poor alignment with human perception. To resolve these limitations, we introduce GenArena, a unified evaluation framework that leverages a pairwise comparison paradigm to ensure stable and human-aligned evaluation. Crucially, our experiments uncover a transformative finding that simply adopting this pairwise protocol enables off-the-shelf open-source models to outperform top-tier proprietary models. Notably, our method boosts evaluation accuracy by over 20% and achieves a Spearman correlation of 0.86 with the authoritative LMArena leaderboard, drastically surpassing the 0.36 correlation of pointwise methods. Based on GenArena, we benchmark state-of-the-art visual generation models across diverse tasks, providing the community with a rigorous and automated evaluation standard for visual generation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06013
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
Li, Ruihang
Qu, Leigang
Zhang, Jingxu
Gui, Dongnan
Xu, Mengde
Zhang, Xiaosong
Hu, Han
Wang, Wenjie
Wang, Jiaqi
Computer Vision and Pattern Recognition
Artificial Intelligence
The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this work, we systematically investigate the reliability of the prevailing absolute pointwise scoring standard, across a wide spectrum of visual generation tasks. Our analysis reveals that this paradigm is limited due to stochastic inconsistency and poor alignment with human perception. To resolve these limitations, we introduce GenArena, a unified evaluation framework that leverages a pairwise comparison paradigm to ensure stable and human-aligned evaluation. Crucially, our experiments uncover a transformative finding that simply adopting this pairwise protocol enables off-the-shelf open-source models to outperform top-tier proprietary models. Notably, our method boosts evaluation accuracy by over 20% and achieves a Spearman correlation of 0.86 with the authoritative LMArena leaderboard, drastically surpassing the 0.36 correlation of pointwise methods. Based on GenArena, we benchmark state-of-the-art visual generation models across diverse tasks, providing the community with a rigorous and automated evaluation standard for visual generation.
title GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.06013