VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915834742439936 |
|---|---|
| author | Park, Kyoungjun Yang, Yifan Yi, Juheon Zheng, Shicheng Shen, Yifei Han, Dongqi Shan, Caihua Muaz, Muhammad Qiu, Lili |
| author_facet | Park, Kyoungjun Yang, Yifan Yi, Juheon Zheng, Shicheng Shen, Yifei Han, Dongqi Shan, Caihua Muaz, Muhammad Qiu, Lili |
| contents | The rapid proliferation of AI-generated video necessitates robust detection tools that offer both high accuracy and human-interpretable explanations. While existing MLLM-based detectors rely on supervised fine-tuning (SFT) or direct preference optimization (DPO), these methods are often bottlenecked by static, pre-labeled datasets that fail to capture the evolving, multi-step physical inconsistencies of modern generative models. To bridge this gap, we introduce VidGuard-R1, the first video authenticity detector to utilize group relative policy optimization (GRPO). Moving beyond passive preference matching, VidGuard-R1 employs a reinforcement learning framework that encourages the model to explore and rank multiple reasoning paths. By introducing specialized reward models for temporal stability and diffusion-aware complexity, we incentivize the model to discover 'physics-grounded' artifacts. Our contributions include: (1) a curated dataset of 140,000 challenging real/fake video pairs; (2) a GRPO-based training paradigm that achieves state-of-the-art zero-shot performance; and (3) a reasoning-first architecture that provides precise, verifiable rationales for its forensic judgments. Project website: https://vidguard-r1.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_02282 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL Park, Kyoungjun Yang, Yifan Yi, Juheon Zheng, Shicheng Shen, Yifei Han, Dongqi Shan, Caihua Muaz, Muhammad Qiu, Lili Computer Vision and Pattern Recognition Machine Learning The rapid proliferation of AI-generated video necessitates robust detection tools that offer both high accuracy and human-interpretable explanations. While existing MLLM-based detectors rely on supervised fine-tuning (SFT) or direct preference optimization (DPO), these methods are often bottlenecked by static, pre-labeled datasets that fail to capture the evolving, multi-step physical inconsistencies of modern generative models. To bridge this gap, we introduce VidGuard-R1, the first video authenticity detector to utilize group relative policy optimization (GRPO). Moving beyond passive preference matching, VidGuard-R1 employs a reinforcement learning framework that encourages the model to explore and rank multiple reasoning paths. By introducing specialized reward models for temporal stability and diffusion-aware complexity, we incentivize the model to discover 'physics-grounded' artifacts. Our contributions include: (1) a curated dataset of 140,000 challenging real/fake video pairs; (2) a GRPO-based training paradigm that achieves state-of-the-art zero-shot performance; and (3) a reasoning-first architecture that provides precise, verifiable rationales for its forensic judgments. Project website: https://vidguard-r1.github.io/. |
| title | VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2510.02282 |