VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Kyoungjun, Yang, Yifan, Yi, Juheon, Zheng, Shicheng, Shen, Yifei, Han, Dongqi, Shan, Caihua, Muaz, Muhammad, Qiu, Lili
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915834742439936
author Park, Kyoungjun
Yang, Yifan
Yi, Juheon
Zheng, Shicheng
Shen, Yifei
Han, Dongqi
Shan, Caihua
Muaz, Muhammad
Qiu, Lili
author_facet Park, Kyoungjun
Yang, Yifan
Yi, Juheon
Zheng, Shicheng
Shen, Yifei
Han, Dongqi
Shan, Caihua
Muaz, Muhammad
Qiu, Lili
contents The rapid proliferation of AI-generated video necessitates robust detection tools that offer both high accuracy and human-interpretable explanations. While existing MLLM-based detectors rely on supervised fine-tuning (SFT) or direct preference optimization (DPO), these methods are often bottlenecked by static, pre-labeled datasets that fail to capture the evolving, multi-step physical inconsistencies of modern generative models. To bridge this gap, we introduce VidGuard-R1, the first video authenticity detector to utilize group relative policy optimization (GRPO). Moving beyond passive preference matching, VidGuard-R1 employs a reinforcement learning framework that encourages the model to explore and rank multiple reasoning paths. By introducing specialized reward models for temporal stability and diffusion-aware complexity, we incentivize the model to discover 'physics-grounded' artifacts. Our contributions include: (1) a curated dataset of 140,000 challenging real/fake video pairs; (2) a GRPO-based training paradigm that achieves state-of-the-art zero-shot performance; and (3) a reasoning-first architecture that provides precise, verifiable rationales for its forensic judgments. Project website: https://vidguard-r1.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
Park, Kyoungjun
Yang, Yifan
Yi, Juheon
Zheng, Shicheng
Shen, Yifei
Han, Dongqi
Shan, Caihua
Muaz, Muhammad
Qiu, Lili
Computer Vision and Pattern Recognition
Machine Learning
The rapid proliferation of AI-generated video necessitates robust detection tools that offer both high accuracy and human-interpretable explanations. While existing MLLM-based detectors rely on supervised fine-tuning (SFT) or direct preference optimization (DPO), these methods are often bottlenecked by static, pre-labeled datasets that fail to capture the evolving, multi-step physical inconsistencies of modern generative models. To bridge this gap, we introduce VidGuard-R1, the first video authenticity detector to utilize group relative policy optimization (GRPO). Moving beyond passive preference matching, VidGuard-R1 employs a reinforcement learning framework that encourages the model to explore and rank multiple reasoning paths. By introducing specialized reward models for temporal stability and diffusion-aware complexity, we incentivize the model to discover 'physics-grounded' artifacts. Our contributions include: (1) a curated dataset of 140,000 challenging real/fake video pairs; (2) a GRPO-based training paradigm that achieves state-of-the-art zero-shot performance; and (3) a reasoning-first architecture that provides precise, verifiable rationales for its forensic judgments. Project website: https://vidguard-r1.github.io/.
title VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.02282