Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.23452 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915261283565568 |
|---|---|
| author | Yang, Yuhang Fan, Ke Sun, Shangkun Li, Hongxiang Zeng, Ailing Han, FeiLin Zhai, Wei Liu, Wei Cao, Yang Zha, Zheng-Jun |
| author_facet | Yang, Yuhang Fan, Ke Sun, Shangkun Li, Hongxiang Zeng, Ailing Han, FeiLin Zhai, Wei Liu, Wei Cao, Yang Zha, Zheng-Jun |
| contents | The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators struggling with Out-of-Distribution (OOD) cases, and misalignment between computed metrics and human preferences. To bridge the gap, we propose VideoGen-Eval, an agent evaluation system that integrates LLM-based content structuring, MLLM-based content judgment, and patch tools designed for temporal-dense dimensions, to achieve a dynamic, flexible, and expandable video generation evaluation. Additionally, we introduce a video generation benchmark to evaluate existing cutting-edge models and verify the effectiveness of our evaluation system. It comprises 700 structured, content-rich prompts (both T2V and I2V) and over 12,000 videos generated by 20+ models, among them, 8 cutting-edge models are selected as quantitative evaluation for the agent and human. Extensive experiments validate that our proposed agent-based evaluation system demonstrates strong alignment with human preferences and reliably completes the evaluation, as well as the diversity and richness of the benchmark. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_23452 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VideoGen-Eval: Agent-based System for Video Generation Evaluation Yang, Yuhang Fan, Ke Sun, Shangkun Li, Hongxiang Zeng, Ailing Han, FeiLin Zhai, Wei Liu, Wei Cao, Yang Zha, Zheng-Jun Computer Vision and Pattern Recognition The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators struggling with Out-of-Distribution (OOD) cases, and misalignment between computed metrics and human preferences. To bridge the gap, we propose VideoGen-Eval, an agent evaluation system that integrates LLM-based content structuring, MLLM-based content judgment, and patch tools designed for temporal-dense dimensions, to achieve a dynamic, flexible, and expandable video generation evaluation. Additionally, we introduce a video generation benchmark to evaluate existing cutting-edge models and verify the effectiveness of our evaluation system. It comprises 700 structured, content-rich prompts (both T2V and I2V) and over 12,000 videos generated by 20+ models, among them, 8 cutting-edge models are selected as quantitative evaluation for the agent and human. Extensive experiments validate that our proposed agent-based evaluation system demonstrates strong alignment with human preferences and reliably completes the evaluation, as well as the diversity and richness of the benchmark. |
| title | VideoGen-Eval: Agent-based System for Video Generation Evaluation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.23452 |