VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Hao, Lan, Jun, Shi, Senyuan, Tan, Zichang, Yu, Zijian, Zhu, Huijia, Wang, Weiqiang, Wan, Jun, Lei, Zhen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908823022731264
author Tan, Hao
Lan, Jun
Shi, Senyuan
Tan, Zichang
Yu, Zijian
Zhu, Huijia
Wang, Weiqiang
Wan, Jun
Lei, Zhen
author_facet Tan, Hao
Lan, Jun
Shi, Senyuan
Tan, Zichang
Yu, Zijian
Zhu, Huijia
Wang, Weiqiang
Wan, Jun
Lei, Zhen
contents The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce VideoVeritas, a framework that integrates fine-grained perception and fact-based reasoning. We observe that while current multi-modal large language models (MLLMs) exhibit strong reasoning capacity, their granular perception ability remains limited. To mitigate this, we introduce Joint Preference Alignment and Perception Pretext Reinforcement Learning (PPRL). Specifically, rather than directly optimizing for detection task, we adopt general spatiotemporal grounding and self-supervised object counting in the RL stage, enhancing detection performance with simple perception pretext tasks. To facilitate robust evaluation, we further introduce MintVid, a light yet high-quality dataset containing 3K videos from 9 state-of-the-art generators, along with a real-world collected subset that has factual errors in content. Experimental results demonstrate that existing methods tend to bias towards either superficial reasoning or mechanical analysis, while VideoVeritas achieves more balanced performance across diverse benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08828
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning
Tan, Hao
Lan, Jun
Shi, Senyuan
Tan, Zichang
Yu, Zijian
Zhu, Huijia
Wang, Weiqiang
Wan, Jun
Lei, Zhen
Computer Vision and Pattern Recognition
The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce VideoVeritas, a framework that integrates fine-grained perception and fact-based reasoning. We observe that while current multi-modal large language models (MLLMs) exhibit strong reasoning capacity, their granular perception ability remains limited. To mitigate this, we introduce Joint Preference Alignment and Perception Pretext Reinforcement Learning (PPRL). Specifically, rather than directly optimizing for detection task, we adopt general spatiotemporal grounding and self-supervised object counting in the RL stage, enhancing detection performance with simple perception pretext tasks. To facilitate robust evaluation, we further introduce MintVid, a light yet high-quality dataset containing 3K videos from 9 state-of-the-art generators, along with a real-world collected subset that has factual errors in content. Experimental results demonstrate that existing methods tend to bias towards either superficial reasoning or mechanical analysis, while VideoVeritas achieves more balanced performance across diverse benchmarks.
title VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.08828