Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Huaxin, Xu, Xiaohao, Wang, Xiang, Zuo, Jialong, Han, Chuchu, Huang, Xiaonan, Gao, Changxin, Wang, Yuehuan, Sang, Nong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909233924014080
author Zhang, Huaxin
Xu, Xiaohao
Wang, Xiang
Zuo, Jialong
Han, Chuchu
Huang, Xiaonan
Gao, Changxin
Wang, Yuehuan
Sang, Nong
author_facet Zhang, Huaxin
Xu, Xiaohao
Wang, Xiang
Zuo, Jialong
Han, Chuchu
Huang, Xiaonan
Gao, Changxin
Wang, Yuehuan
Sang, Nong
contents Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD, a novel framework that leverages precise temporal supervision and rich multimodal instructions to enable accurate anomaly localization and comprehensive explanations. Firstly, towards unbiased and explainable VAD system, we construct the first large-scale multimodal VAD instruction-tuning benchmark, i.e., VAD-Instruct50k. This dataset is created using a carefully designed semi-automatic labeling paradigm. Efficient single-frame annotations are applied to the collected untrimmed videos, which are then synthesized into high-quality analyses of both abnormal and normal video clips using a robust off-the-shelf video captioner and a large language model (LLM). Building upon the VAD-Instruct50k dataset, we develop a customized solution for interpretable video anomaly detection. We train a lightweight temporal sampler to select frames with high anomaly response and fine-tune a multimodal large language model (LLM) to generate explanatory content. Extensive experimental results validate the generality and interpretability of the proposed Holmes-VAD, establishing it as a novel interpretable technique for real-world video anomaly analysis. To support the community, our benchmark and model will be publicly available at https://holmesvad.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12235
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
Zhang, Huaxin
Xu, Xiaohao
Wang, Xiang
Zuo, Jialong
Han, Chuchu
Huang, Xiaonan
Gao, Changxin
Wang, Yuehuan
Sang, Nong
Computer Vision and Pattern Recognition
Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD, a novel framework that leverages precise temporal supervision and rich multimodal instructions to enable accurate anomaly localization and comprehensive explanations. Firstly, towards unbiased and explainable VAD system, we construct the first large-scale multimodal VAD instruction-tuning benchmark, i.e., VAD-Instruct50k. This dataset is created using a carefully designed semi-automatic labeling paradigm. Efficient single-frame annotations are applied to the collected untrimmed videos, which are then synthesized into high-quality analyses of both abnormal and normal video clips using a robust off-the-shelf video captioner and a large language model (LLM). Building upon the VAD-Instruct50k dataset, we develop a customized solution for interpretable video anomaly detection. We train a lightweight temporal sampler to select frames with high anomaly response and fine-tune a multimodal large language model (LLM) to generate explanatory content. Extensive experimental results validate the generality and interpretability of the proposed Holmes-VAD, establishing it as a novel interpretable technique for real-world video anomaly analysis. To support the community, our benchmark and model will be publicly available at https://holmesvad.github.io.
title Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.12235