VACT: A Video Automatic Causal Testing System and a Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Haotong, Zheng, Qingyuan, Gao, Yunjian, Yang, Yongkun, He, Yangbo, Lin, Zhouchen, Zhang, Muhan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910914903539712
author Yang, Haotong
Zheng, Qingyuan
Gao, Yunjian
Yang, Yongkun
He, Yangbo
Lin, Zhouchen
Zhang, Muhan
author_facet Yang, Haotong
Zheng, Qingyuan
Gao, Yunjian
Yang, Yongkun
He, Yangbo
Lin, Zhouchen
Zhang, Muhan
contents With the rapid advancement of text-conditioned Video Generation Models (VGMs), the quality of generated videos has significantly improved, bringing these models closer to functioning as ``*world simulators*'' and making real-world-level video generation more accessible and cost-effective. However, the generated videos often contain factual inaccuracies and lack understanding of fundamental physical laws. While some previous studies have highlighted this issue in limited domains through manual analysis, a comprehensive solution has not yet been established, primarily due to the absence of a generalized, automated approach for modeling and assessing the causal reasoning of these models across diverse scenarios. To address this gap, we propose VACT: an **automated** framework for modeling, evaluating, and measuring the causal understanding of VGMs in real-world scenarios. By combining causal analysis techniques with a carefully designed large language model assistant, our system can assess the causal behavior of models in various contexts without human annotation, which offers strong generalization and scalability. Additionally, we introduce multi-level causal evaluation metrics to provide a detailed analysis of the causal performance of VGMs. As a demonstration, we use our framework to benchmark several prevailing VGMs, offering insight into their causal reasoning capabilities. Our work lays the foundation for systematically addressing the causal understanding deficiencies in VGMs and contributes to advancing their reliability and real-world applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06163
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VACT: A Video Automatic Causal Testing System and a Benchmark
Yang, Haotong
Zheng, Qingyuan
Gao, Yunjian
Yang, Yongkun
He, Yangbo
Lin, Zhouchen
Zhang, Muhan
Artificial Intelligence
Computer Vision and Pattern Recognition
Applications
With the rapid advancement of text-conditioned Video Generation Models (VGMs), the quality of generated videos has significantly improved, bringing these models closer to functioning as ``*world simulators*'' and making real-world-level video generation more accessible and cost-effective. However, the generated videos often contain factual inaccuracies and lack understanding of fundamental physical laws. While some previous studies have highlighted this issue in limited domains through manual analysis, a comprehensive solution has not yet been established, primarily due to the absence of a generalized, automated approach for modeling and assessing the causal reasoning of these models across diverse scenarios. To address this gap, we propose VACT: an **automated** framework for modeling, evaluating, and measuring the causal understanding of VGMs in real-world scenarios. By combining causal analysis techniques with a carefully designed large language model assistant, our system can assess the causal behavior of models in various contexts without human annotation, which offers strong generalization and scalability. Additionally, we introduce multi-level causal evaluation metrics to provide a detailed analysis of the causal performance of VGMs. As a demonstration, we use our framework to benchmark several prevailing VGMs, offering insight into their causal reasoning capabilities. Our work lays the foundation for systematically addressing the causal understanding deficiencies in VGMs and contributes to advancing their reliability and real-world applicability.
title VACT: A Video Automatic Causal Testing System and a Benchmark
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Applications
url https://arxiv.org/abs/2503.06163