GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Jenna, Silva, Maria, Sangkloy, Patsorn, Chen, Kenneth, Williams, Niall, Sun, Qi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914031515729920
author Kang, Jenna
Silva, Maria
Sangkloy, Patsorn
Chen, Kenneth
Williams, Niall
Sun, Qi
author_facet Kang, Jenna
Silva, Maria
Sangkloy, Patsorn
Chen, Kenneth
Williams, Niall
Sun, Qi
contents Recent advances in probabilistic generative models have extended capabilities from static image synthesis to text-driven video generation. However, the inherent randomness of their generation process can lead to unpredictable artifacts, such as impossible physics and temporal inconsistency. Progress in addressing these challenges requires systematic benchmarks, yet existing datasets primarily focus on generative images due to the unique spatio-temporal complexities of videos. To bridge this gap, we introduce GeneVA, a large-scale artifact dataset with rich human annotations that focuses on spatio-temporal artifacts in videos generated from natural text prompts. We hope GeneVA can enable and assist critical applications, such as benchmarking model performance and improving generative video quality.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
Kang, Jenna
Silva, Maria
Sangkloy, Patsorn
Chen, Kenneth
Williams, Niall
Sun, Qi
Computer Vision and Pattern Recognition
Recent advances in probabilistic generative models have extended capabilities from static image synthesis to text-driven video generation. However, the inherent randomness of their generation process can lead to unpredictable artifacts, such as impossible physics and temporal inconsistency. Progress in addressing these challenges requires systematic benchmarks, yet existing datasets primarily focus on generative images due to the unique spatio-temporal complexities of videos. To bridge this gap, we introduce GeneVA, a large-scale artifact dataset with rich human annotations that focuses on spatio-temporal artifacts in videos generated from natural text prompts. We hope GeneVA can enable and assist critical applications, such as benchmarking model performance and improving generative video quality.
title GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.08818