VideoScore2: Think before You Score in Generative Video Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Xuan, Jiang, Dongfu, Nie, Ping, Liu, Minghao, Jiang, Zhengxuan, Su, Mingyi, Ma, Wentao, Lin, Junru, Ye, Chun, Lu, Yi, Wu, Keming, Schneider, Benjamin, Do, Quy Duc, Li, Zhuofeng, Jia, Yiming, Zhang, Yuxuan, Cheng, Guo, Wang, Haozhe, Zhou, Wangchunshu, Lin, Qunshu, Zhang, Yuanxing, Zhang, Ge, Huang, Wenhao, Chen, Wenhu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912609446395904
author He, Xuan
Jiang, Dongfu
Nie, Ping
Liu, Minghao
Jiang, Zhengxuan
Su, Mingyi
Ma, Wentao
Lin, Junru
Ye, Chun
Lu, Yi
Wu, Keming
Schneider, Benjamin
Do, Quy Duc
Li, Zhuofeng
Jia, Yiming
Zhang, Yuxuan
Cheng, Guo
Wang, Haozhe
Zhou, Wangchunshu
Lin, Qunshu
Zhang, Yuanxing
Zhang, Ge
Huang, Wenhao
Chen, Wenhu
author_facet He, Xuan
Jiang, Dongfu
Nie, Ping
Liu, Minghao
Jiang, Zhengxuan
Su, Mingyi
Ma, Wentao
Lin, Junru
Ye, Chun
Lu, Yi
Wu, Keming
Schneider, Benjamin
Do, Quy Duc
Li, Zhuofeng
Jia, Yiming
Zhang, Yuxuan
Cheng, Guo
Wang, Haozhe
Zhou, Wangchunshu
Lin, Qunshu
Zhang, Yuanxing
Zhang, Ge
Huang, Wenhao
Chen, Wenhu
contents Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales. Our model is trained on a large-scale dataset VideoFeedback2 containing 27,168 human-annotated videos with both scores and reasoning traces across three dimensions, using a two-stage pipeline of supervised fine-tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in-domain benchmark VideoScore-Bench-v2 and 50.37 (+4.32) average performance across four out-of-domain benchmarks (VideoGenReward-Bench, VideoPhy2, etc), while providing interpretable assessments that bridge the gap between evaluation and controllable generation through effective reward modeling for Best-of-N sampling. Project Page: https://tiger-ai-lab.github.io/VideoScore2/
format Preprint
id arxiv_https___arxiv_org_abs_2509_22799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoScore2: Think before You Score in Generative Video Evaluation
He, Xuan
Jiang, Dongfu
Nie, Ping
Liu, Minghao
Jiang, Zhengxuan
Su, Mingyi
Ma, Wentao
Lin, Junru
Ye, Chun
Lu, Yi
Wu, Keming
Schneider, Benjamin
Do, Quy Duc
Li, Zhuofeng
Jia, Yiming
Zhang, Yuxuan
Cheng, Guo
Wang, Haozhe
Zhou, Wangchunshu
Lin, Qunshu
Zhang, Yuanxing
Zhang, Ge
Huang, Wenhao
Chen, Wenhu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales. Our model is trained on a large-scale dataset VideoFeedback2 containing 27,168 human-annotated videos with both scores and reasoning traces across three dimensions, using a two-stage pipeline of supervised fine-tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in-domain benchmark VideoScore-Bench-v2 and 50.37 (+4.32) average performance across four out-of-domain benchmarks (VideoGenReward-Bench, VideoPhy2, etc), while providing interpretable assessments that bridge the gap between evaluation and controllable generation through effective reward modeling for Best-of-N sampling. Project Page: https://tiger-ai-lab.github.io/VideoScore2/
title VideoScore2: Think before You Score in Generative Video Evaluation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.22799