Q-Save: Towards Scoring and Attribution for Generated Video Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xiele, Zhang, Zicheng, Chen, Mingtao, Liu, Yixian, Liu, Yiming, Wang, Shushi, Hu, Zhichao, Liu, Yuhong, Zhai, Guangtao, Liu, Xiaohong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914355065389056
author Wu, Xiele
Zhang, Zicheng
Chen, Mingtao
Liu, Yixian
Liu, Yiming
Wang, Shushi
Hu, Zhichao
Liu, Yuhong
Zhai, Guangtao
Liu, Xiaohong
author_facet Wu, Xiele
Zhang, Zicheng
Chen, Mingtao
Liu, Yixian
Liu, Yiming
Wang, Shushi
Hu, Zhichao
Liu, Yuhong
Zhai, Guangtao
Liu, Xiaohong
contents Evaluating AI-generated video (AIGV) quality hinges on three crucial dimensions: visual quality, dynamic quality, and text-video alignment. While numerous evaluation datasets and algorithms have been proposed, existing approaches are constrained by two limitations: the absence of systematic definitions for evaluation dimensions, and the isolated treatment of the three dimensions in separate models. Therefore, we introduce Q-Save, a holistic benchmark dataset and unified evaluation model for AIGV quality assessment. The Q-Save dataset contains nearly 10,000 video samples, each annotated with Mean Opinion Scores (MOS) and fine-grained attribution explanations across the three core dimensions. Leveraging this attribution-annotated dataset, we train the proposed Q-Save model, which adopts the SlowFast framework to balance accuracy and efficiency, and employs a three-stage training strategy with Chain-of-Thought (COT) formatted data: Supervised Fine-Tuning (SFT), Grouped Relative Policy Optimization (GRPO), and a final SFT round for stability, to jointly perform quality scoring and attribution generation. Experimental results demonstrate that Q-Save achieves superior performance in AIGV quality prediction while providing interpretable justifications. Code and dataset will be released upon publication.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18825
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Q-Save: Towards Scoring and Attribution for Generated Video Evaluation
Wu, Xiele
Zhang, Zicheng
Chen, Mingtao
Liu, Yixian
Liu, Yiming
Wang, Shushi
Hu, Zhichao
Liu, Yuhong
Zhai, Guangtao
Liu, Xiaohong
Computer Vision and Pattern Recognition
Evaluating AI-generated video (AIGV) quality hinges on three crucial dimensions: visual quality, dynamic quality, and text-video alignment. While numerous evaluation datasets and algorithms have been proposed, existing approaches are constrained by two limitations: the absence of systematic definitions for evaluation dimensions, and the isolated treatment of the three dimensions in separate models. Therefore, we introduce Q-Save, a holistic benchmark dataset and unified evaluation model for AIGV quality assessment. The Q-Save dataset contains nearly 10,000 video samples, each annotated with Mean Opinion Scores (MOS) and fine-grained attribution explanations across the three core dimensions. Leveraging this attribution-annotated dataset, we train the proposed Q-Save model, which adopts the SlowFast framework to balance accuracy and efficiency, and employs a three-stage training strategy with Chain-of-Thought (COT) formatted data: Supervised Fine-Tuning (SFT), Grouped Relative Policy Optimization (GRPO), and a final SFT round for stability, to jointly perform quality scoring and attribution generation. Experimental results demonstrate that Q-Save achieves superior performance in AIGV quality prediction while providing interpretable justifications. Code and dataset will be released upon publication.
title Q-Save: Towards Scoring and Attribution for Generated Video Evaluation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18825