Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hayun, Omer Ben, Betser, Roy, Levi, Meir Yossef, Kassel, Levi, Gilboa, Guy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908895751962624
author Hayun, Omer Ben
Betser, Roy
Levi, Meir Yossef
Kassel, Levi
Gilboa, Guy
author_facet Hayun, Omer Ben
Betser, Roy
Levi, Meir Yossef
Kassel, Levi
Gilboa, Guy
contents Following major advances in text and image generation, the video domain has surged, producing highly realistic and controllable sequences. Along with this progress, these models also raise serious concerns about misinformation, making reliable detection of synthetic videos increasingly crucial. Image-based detectors are fundamentally limited because they operate per frame and ignore temporal dynamics, while supervised video detectors generalize poorly to unseen generators, a critical drawback given the rapid emergence of new models. These challenges motivate zero-shot approaches, which avoid synthetic data and instead score content against real-data statistics, enabling training-free, model-agnostic detection. We introduce STALL, a simple, training-free, theoretically justified detector that provides likelihood-based scoring for videos, jointly modeling spatial and temporal evidence within a probabilistic framework. We evaluate STALL on two public benchmarks and introduce ComGenVid, a new benchmark with state-of-the-art generative models. STALL consistently outperforms prior image- and video-based baselines. Code and data are available at https://omerbenhayun.github.io/stall-video.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15026
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
Hayun, Omer Ben
Betser, Roy
Levi, Meir Yossef
Kassel, Levi
Gilboa, Guy
Computer Vision and Pattern Recognition
Machine Learning
Following major advances in text and image generation, the video domain has surged, producing highly realistic and controllable sequences. Along with this progress, these models also raise serious concerns about misinformation, making reliable detection of synthetic videos increasingly crucial. Image-based detectors are fundamentally limited because they operate per frame and ignore temporal dynamics, while supervised video detectors generalize poorly to unseen generators, a critical drawback given the rapid emergence of new models. These challenges motivate zero-shot approaches, which avoid synthetic data and instead score content against real-data statistics, enabling training-free, model-agnostic detection. We introduce STALL, a simple, training-free, theoretically justified detector that provides likelihood-based scoring for videos, jointly modeling spatial and temporal evidence within a probabilistic framework. We evaluate STALL on two public benchmarks and introduce ComGenVid, a new benchmark with state-of-the-art generative models. STALL consistently outperforms prior image- and video-based baselines. Code and data are available at https://omerbenhayun.github.io/stall-video.
title Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.15026