ProDisc-VAD: An Efficient System for Weakly-Supervised Anomaly Detection in Video Surveillance Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Tao, Yu, Qi, Dong, Xinru, Li, Shiyu, Liu, Yue, Jiang, Jinlong, Shu, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912487260028928
author Zhu, Tao
Yu, Qi
Dong, Xinru
Li, Shiyu
Liu, Yue
Jiang, Jinlong
Shu, Lei
author_facet Zhu, Tao
Yu, Qi
Dong, Xinru
Li, Shiyu
Liu, Yue
Jiang, Jinlong
Shu, Lei
contents Weakly-supervised video anomaly detection (WS-VAD) using Multiple Instance Learning (MIL) suffers from label ambiguity, hindering discriminative feature learning. We propose ProDisc-VAD, an efficient framework tackling this via two synergistic components. The Prototype Interaction Layer (PIL) provides controlled normality modeling using a small set of learnable prototypes, establishing a robust baseline without being overwhelmed by dominant normal data. The Pseudo-Instance Discriminative Enhancement (PIDE) loss boosts separability by applying targeted contrastive learning exclusively to the most reliable extreme-scoring instances (highest/lowest scores). ProDisc-VAD achieves strong AUCs (97.98% ShanghaiTech, 87.12% UCF-Crime) using only 0.4M parameters, over 800x fewer than recent ViT-based methods like VadCLIP. Code is available at https://github.com/modadundun/ProDisc-VAD.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProDisc-VAD: An Efficient System for Weakly-Supervised Anomaly Detection in Video Surveillance Applications
Zhu, Tao
Yu, Qi
Dong, Xinru
Li, Shiyu
Liu, Yue
Jiang, Jinlong
Shu, Lei
Computer Vision and Pattern Recognition
Weakly-supervised video anomaly detection (WS-VAD) using Multiple Instance Learning (MIL) suffers from label ambiguity, hindering discriminative feature learning. We propose ProDisc-VAD, an efficient framework tackling this via two synergistic components. The Prototype Interaction Layer (PIL) provides controlled normality modeling using a small set of learnable prototypes, establishing a robust baseline without being overwhelmed by dominant normal data. The Pseudo-Instance Discriminative Enhancement (PIDE) loss boosts separability by applying targeted contrastive learning exclusively to the most reliable extreme-scoring instances (highest/lowest scores). ProDisc-VAD achieves strong AUCs (97.98% ShanghaiTech, 87.12% UCF-Crime) using only 0.4M parameters, over 800x fewer than recent ViT-based methods like VadCLIP. Code is available at https://github.com/modadundun/ProDisc-VAD.
title ProDisc-VAD: An Efficient System for Weakly-Supervised Anomaly Detection in Video Surveillance Applications
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.02179