ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Kaisi, Lai, Zhengfeng, Sun, Yuchong, Zhang, Peng, Liu, Wei, Liu, Kieran, Cao, Meng, Song, Ruihua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912540355723264
author Guan, Kaisi
Lai, Zhengfeng
Sun, Yuchong
Zhang, Peng
Liu, Wei
Liu, Kieran
Cao, Meng
Song, Ruihua
author_facet Guan, Kaisi
Lai, Zhengfeng
Sun, Yuchong
Zhang, Peng
Liu, Wei
Liu, Kieran
Cao, Meng
Song, Ruihua
contents Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without fine-grained alignment details, failing to align with human preference. To address this limitation, we propose ETVA, a novel Evaluation method of Text-to-Video Alignment via fine-grained question generation and answering. First, a multi-agent system parses prompts into semantic scene graphs to generate atomic questions. Then we design a knowledge-augmented multi-stage reasoning framework for question answering, where an auxiliary LLM first retrieves relevant common-sense knowledge (e.g., physical laws), and then video LLM answers the generated questions through a multi-stage reasoning mechanism. Extensive experiments demonstrate that ETVA achieves a Spearman's correlation coefficient of 58.47, showing a much higher correlation with human judgment than existing metrics which attain only 31.0. We also construct a comprehensive benchmark specifically designed for text-to-video alignment evaluation, featuring 2k diverse prompts and 12k atomic questions spanning 10 categories. Through a systematic evaluation of 15 existing text-to-video models, we identify their key capabilities and limitations, paving the way for next-generation T2V generation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
Guan, Kaisi
Lai, Zhengfeng
Sun, Yuchong
Zhang, Peng
Liu, Wei
Liu, Kieran
Cao, Meng
Song, Ruihua
Computer Vision and Pattern Recognition
Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without fine-grained alignment details, failing to align with human preference. To address this limitation, we propose ETVA, a novel Evaluation method of Text-to-Video Alignment via fine-grained question generation and answering. First, a multi-agent system parses prompts into semantic scene graphs to generate atomic questions. Then we design a knowledge-augmented multi-stage reasoning framework for question answering, where an auxiliary LLM first retrieves relevant common-sense knowledge (e.g., physical laws), and then video LLM answers the generated questions through a multi-stage reasoning mechanism. Extensive experiments demonstrate that ETVA achieves a Spearman's correlation coefficient of 58.47, showing a much higher correlation with human judgment than existing metrics which attain only 31.0. We also construct a comprehensive benchmark specifically designed for text-to-video alignment evaluation, featuring 2k diverse prompts and 12k atomic questions spanning 10 categories. Through a systematic evaluation of 15 existing text-to-video models, we identify their key capabilities and limitations, paving the way for next-generation T2V generation.
title ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.16867