VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Liyun, Chen, Qixiang, Shen, Xi, Cun, Xiaodong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916766503927808
author Zhu, Liyun
Chen, Qixiang
Shen, Xi
Cun, Xiaodong
author_facet Zhu, Liyun
Chen, Qixiang
Shen, Xi
Cun, Xiaodong
contents Video Anomaly Understanding (VAU) is essential for applications such as smart cities, security surveillance, and disaster alert systems, yet remains challenging due to its demand for fine-grained spatio-temporal perception and robust reasoning under ambiguity. Despite advances in anomaly detection, existing methods often lack interpretability and struggle to capture the causal and contextual aspects of abnormal events. This limitation is further compounded by the absence of comprehensive benchmarks for evaluating reasoning ability in anomaly scenarios. To address both challenges, we introduce VAU-R1, a data-efficient framework built upon Multimodal Large Language Models (MLLMs), which enhances anomaly reasoning through Reinforcement Fine-Tuning (RFT). Besides, we propose VAU-Bench, the first Chain-of-Thought benchmark tailored for video anomaly reasoning, featuring multiple-choice QA, detailed rationales, temporal annotations, and descriptive captions. Empirical results show that VAU-R1 significantly improves question answering accuracy, temporal grounding, and reasoning coherence across diverse contexts. Together, our method and benchmark establish a strong foundation for interpretable and reasoning-aware video anomaly understanding. Our code is available at https://github.com/GVCLab/VAU-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23504
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
Zhu, Liyun
Chen, Qixiang
Shen, Xi
Cun, Xiaodong
Computer Vision and Pattern Recognition
Video Anomaly Understanding (VAU) is essential for applications such as smart cities, security surveillance, and disaster alert systems, yet remains challenging due to its demand for fine-grained spatio-temporal perception and robust reasoning under ambiguity. Despite advances in anomaly detection, existing methods often lack interpretability and struggle to capture the causal and contextual aspects of abnormal events. This limitation is further compounded by the absence of comprehensive benchmarks for evaluating reasoning ability in anomaly scenarios. To address both challenges, we introduce VAU-R1, a data-efficient framework built upon Multimodal Large Language Models (MLLMs), which enhances anomaly reasoning through Reinforcement Fine-Tuning (RFT). Besides, we propose VAU-Bench, the first Chain-of-Thought benchmark tailored for video anomaly reasoning, featuring multiple-choice QA, detailed rationales, temporal annotations, and descriptive captions. Empirical results show that VAU-R1 significantly improves question answering accuracy, temporal grounding, and reasoning coherence across diverse contexts. Together, our method and benchmark establish a strong foundation for interpretable and reasoning-aware video anomaly understanding. Our code is available at https://github.com/GVCLab/VAU-R1.
title VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23504