Saved in:
Bibliographic Details
Main Authors: Wei, Zixi, Zhang, Huixuaun, Wan, Xiaojun
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.17311
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916020540669952
author Wei, Zixi
Zhang, Huixuaun
Wan, Xiaojun
author_facet Wei, Zixi
Zhang, Huixuaun
Wan, Xiaojun
contents The remarkable visual fidelity of recent commercial video generative models, such as Sora and Veo, renders robust AI-generated video detection increasingly essential to prevent synthetic content from being indistinguishable from real videos and exploited for disinformation. However, existing detectors often fail due to an over-reliance on increasingly realistic semantic features, neglecting subtle spectral artifacts. In this paper, we propose SpecSem-Net, the first framework to introduce a semantic-guided spectral denoising mechanism specifically for high-fidelity AI-generated video detection. Specifically, we design a spectral module to extract high-frequency features via Fourier-Transform based filtering. Furthermore, to reduce misjudgments arising from spectral noise, we employ a Gated Merging Mechanism to adaptively fuse semantic context, effectively mitigating spectral noise. Additionally, to evaluate detector performance on the latest top-tier generative models, we construct a comprehensive benchmark comprising 5 SOTA commercial generators. Extensive experiments demonstrate that SpecSem-Net outperforms existing methods, achieving accuracies of 87.25% and 95.59% on our benchmark and public datasets, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17311
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection
Wei, Zixi
Zhang, Huixuaun
Wan, Xiaojun
Computer Vision and Pattern Recognition
The remarkable visual fidelity of recent commercial video generative models, such as Sora and Veo, renders robust AI-generated video detection increasingly essential to prevent synthetic content from being indistinguishable from real videos and exploited for disinformation. However, existing detectors often fail due to an over-reliance on increasingly realistic semantic features, neglecting subtle spectral artifacts. In this paper, we propose SpecSem-Net, the first framework to introduce a semantic-guided spectral denoising mechanism specifically for high-fidelity AI-generated video detection. Specifically, we design a spectral module to extract high-frequency features via Fourier-Transform based filtering. Furthermore, to reduce misjudgments arising from spectral noise, we employ a Gated Merging Mechanism to adaptively fuse semantic context, effectively mitigating spectral noise. Additionally, to evaluate detector performance on the latest top-tier generative models, we construct a comprehensive benchmark comprising 5 SOTA commercial generators. Extensive experiments demonstrate that SpecSem-Net outperforms existing methods, achieving accuracies of 87.25% and 95.59% on our benchmark and public datasets, respectively.
title SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17311