UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuanxin, Zhu, Rui, Ren, Shuhuai, Wang, Jiacong, Guo, Haoyuan, Sun, Xu, Jiang, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908771591127040
author Liu, Yuanxin
Zhu, Rui
Ren, Shuhuai
Wang, Jiacong
Guo, Haoyuan
Sun, Xu
Jiang, Lu
author_facet Liu, Yuanxin
Zhu, Rui
Ren, Shuhuai
Wang, Jiacong
Guo, Haoyuan
Sun, Xu
Jiang, Lu
contents With the rapid growth of video generative models (VGMs), it is essential to develop reliable and comprehensive automatic metrics for AI-generated videos (AIGVs). Existing methods either use off-the-shelf models optimized for other tasks or rely on human assessment data to train specialized evaluators. These approaches are constrained to specific evaluation aspects and are difficult to scale with the increasing demands for finer-grained and more comprehensive evaluations. To address this issue, this work investigates the feasibility of using multimodal large language models (MLLMs) as a unified evaluator for AIGVs, leveraging their strong visual perception and language understanding capabilities. To evaluate the performance of automatic metrics in unified AIGV evaluation, we introduce a benchmark called UVE-Bench. UVE-Bench collects videos generated by state-of-the-art VGMs and provides pairwise human preference annotations across 15 evaluation aspects. Using UVE-Bench, we extensively evaluate 18 MLLMs. Our empirical results suggest that while advanced MLLMs (e.g., Qwen2VL-72B and InternVL2.5-78B) still lag behind human evaluators, they demonstrate promising ability in unified AIGV evaluation, significantly surpassing existing specialized evaluation methods. Additionally, we conduct an in-depth analysis of key design choices that impact the performance of MLLM-driven evaluators, offering valuable insights for future research on AIGV evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
Liu, Yuanxin
Zhu, Rui
Ren, Shuhuai
Wang, Jiacong
Guo, Haoyuan
Sun, Xu
Jiang, Lu
Computer Vision and Pattern Recognition
With the rapid growth of video generative models (VGMs), it is essential to develop reliable and comprehensive automatic metrics for AI-generated videos (AIGVs). Existing methods either use off-the-shelf models optimized for other tasks or rely on human assessment data to train specialized evaluators. These approaches are constrained to specific evaluation aspects and are difficult to scale with the increasing demands for finer-grained and more comprehensive evaluations. To address this issue, this work investigates the feasibility of using multimodal large language models (MLLMs) as a unified evaluator for AIGVs, leveraging their strong visual perception and language understanding capabilities. To evaluate the performance of automatic metrics in unified AIGV evaluation, we introduce a benchmark called UVE-Bench. UVE-Bench collects videos generated by state-of-the-art VGMs and provides pairwise human preference annotations across 15 evaluation aspects. Using UVE-Bench, we extensively evaluate 18 MLLMs. Our empirical results suggest that while advanced MLLMs (e.g., Qwen2VL-72B and InternVL2.5-78B) still lag behind human evaluators, they demonstrate promising ability in unified AIGV evaluation, significantly surpassing existing specialized evaluation methods. Additionally, we conduct an in-depth analysis of key design choices that impact the performance of MLLM-driven evaluators, offering valuable insights for future research on AIGV evaluation.
title UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.09949