VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Kyoungjun, Yang, Yifan, Yi, Juheon, Zheng, Shicheng, Shen, Yifei, Han, Dongqi, Shan, Caihua, Muaz, Muhammad, Qiu, Lili |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint Optimization of Handoff and Video Rate in LEO Satellite Networks
by: Park, Kyoungjun, et al.
Published: (2025)
by: Park, Kyoungjun, et al.
Published: (2025)
Diffusion^2: Turning 3D Environments into Radio Frequency Heatmaps
by: Park, Kyoungjun, et al.
Published: (2025)
by: Park, Kyoungjun, et al.
Published: (2025)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
NeuSaver: Neural Adaptive Power Consumption Optimization for Mobile Video Streaming
by: Park, Kyoungjun, et al.
Published: (2021)
by: Park, Kyoungjun, et al.
Published: (2021)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
Towards Graph Foundation Models: Training on Knowledge Graphs Enables Transferability to General Graphs
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
Kuramoto Oscillatory Phase Encoding: Neuro-inspired Synchronization for Improved Learning Efficiency
by: Xiao, Mingqing, et al.
Published: (2026)
by: Xiao, Mingqing, et al.
Published: (2026)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Revisiting the Graph Reasoning Ability of Large Language Models: Case Studies in Translation, Connectivity and Shortest Path
by: Dai, Xinnan, et al.
Published: (2024)
by: Dai, Xinnan, et al.
Published: (2024)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
by: Meng, Desen, et al.
Published: (2025)
by: Meng, Desen, et al.
Published: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
ARS: Adaptive Reasoning Suppression for Efficient Large Reasoning Language Models
by: Zheng, Dongqi
Published: (2025)
by: Zheng, Dongqi
Published: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Mondrian: On-Device High-Performance Video Analytics with Compressive Packed Inference
by: Jeon, Changmin, et al.
Published: (2024)
by: Jeon, Changmin, et al.
Published: (2024)
Resurrecting Label Propagation for Graphs with Heterophily and Label Noise
by: Cheng, Yao, et al.
Published: (2023)
by: Cheng, Yao, et al.
Published: (2023)
PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning
by: Yan, Sikuan, et al.
Published: (2026)
by: Yan, Sikuan, et al.
Published: (2026)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
by: Muaz, Muhammad, et al.
Published: (2024)
by: Muaz, Muhammad, et al.
Published: (2024)
Understanding and Improving Training-free Loss-based Diffusion Guidance
by: Shen, Yifei, et al.
Published: (2024)
by: Shen, Yifei, et al.
Published: (2024)
Enabling Cross-Camera Collaboration for Video Analytics on Distributed Smart Cameras
by: Min, Chulhong, et al.
Published: (2024)
by: Min, Chulhong, et al.
Published: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
What Makes a Good Diffusion Planner for Decision Making?
by: Lu, Haofei, et al.
Published: (2025)
by: Lu, Haofei, et al.
Published: (2025)
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
by: Wen, Zimo, et al.
Published: (2026)
by: Wen, Zimo, et al.
Published: (2026)
Vid2Sid: Videos Can Help Close the Sim2Real Gap
by: Qiu, Kevin, et al.
Published: (2026)
by: Qiu, Kevin, et al.
Published: (2026)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
by: Wen, Youpeng, et al.
Published: (2024)
by: Wen, Youpeng, et al.
Published: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
When Do LLMs Help With Node Classification? A Comprehensive Analysis
by: Wu, Xixi, et al.
Published: (2025)
by: Wu, Xixi, et al.
Published: (2025)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
by: Lai, Yingxin, et al.
Published: (2026)
by: Lai, Yingxin, et al.
Published: (2026)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
Papez: Resource-Efficient Speech Separation with Auditory Working Memory
by: Oh, Hyunseok, et al.
Published: (2024)
by: Oh, Hyunseok, et al.
Published: (2024)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
by: Liang, Lili, et al.
Published: (2024)
by: Liang, Lili, et al.
Published: (2024)
Similar Items
-
Joint Optimization of Handoff and Video Rate in LEO Satellite Networks
by: Park, Kyoungjun, et al.
Published: (2025) -
Diffusion^2: Turning 3D Environments into Radio Frequency Heatmaps
by: Park, Kyoungjun, et al.
Published: (2025) -
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024) -
NeuSaver: Neural Adaptive Power Consumption Optimization for Mobile Video Streaming
by: Park, Kyoungjun, et al.
Published: (2021) -
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)