Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Huaxin, Xu, Xiaohao, Wang, Xiang, Zuo, Jialong, Han, Chuchu, Huang, Xiaonan, Gao, Changxin, Wang, Yuehuan, Sang, Nong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly Detection
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment
by: Yin, Wenti, et al.
Published: (2025)
by: Yin, Wenti, et al.
Published: (2025)
HR-Pro: Point-supervised Temporal Action Localization via Hierarchical Reliability Propagation
by: Zhang, Huaxin, et al.
Published: (2023)
by: Zhang, Huaxin, et al.
Published: (2023)
Spatial Cascaded Clustering and Weighted Memory for Unsupervised Person Re-identification
by: Hong, Jiahao, et al.
Published: (2024)
by: Hong, Jiahao, et al.
Published: (2024)
GV-VAD : Exploring Video Generation for Weakly-Supervised Video Anomaly Detection
by: Cai, Suhang, et al.
Published: (2025)
by: Cai, Suhang, et al.
Published: (2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
Cross-video Identity Correlating for Person Re-identification Pre-training
by: Zuo, Jialong, et al.
Published: (2024)
by: Zuo, Jialong, et al.
Published: (2024)
Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training Acceleration
by: Wu, Dongyue, et al.
Published: (2025)
by: Wu, Dongyue, et al.
Published: (2025)
ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
UFineBench: Towards Text-based Person Retrieval with Ultra-fine Granularity
by: Zuo, Jialong, et al.
Published: (2023)
by: Zuo, Jialong, et al.
Published: (2023)
Replace Anyone in Videos
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
Taming Consistency Distillation for Accelerated Human Image Animation
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
ComplexVAD: Detecting Interaction Anomalies in Video
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
PLIP: Language-Image Pre-training for Person Representation Learning
by: Zuo, Jialong, et al.
Published: (2023)
by: Zuo, Jialong, et al.
Published: (2023)
DFIMat: Decoupled Flexible Interactive Matting in Multi-Person Scenarios
by: Jiao, Siyi, et al.
Published: (2024)
by: Jiao, Siyi, et al.
Published: (2024)
EventVAD: Training-Free Event-Aware Video Anomaly Detection
by: Shao, Yihua, et al.
Published: (2025)
by: Shao, Yihua, et al.
Published: (2025)
DMPT: Decoupled Modality-aware Prompt Tuning for Multi-modal Object Re-identification
by: Lin, Minghui, et al.
Published: (2025)
by: Lin, Minghui, et al.
Published: (2025)
Object-Aware Video Matting with Cross-Frame Guidance
by: Zhang, Huayu, et al.
Published: (2025)
by: Zhang, Huayu, et al.
Published: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
DUAL-VAD: Dual Benchmarks and Anomaly-Focused Sampling for Video Anomaly Detection
by: Jung, Seoik, et al.
Published: (2025)
by: Jung, Seoik, et al.
Published: (2025)
HyCoVAD: A Hybrid SSL-LLM Model for Complex Video Anomaly Detection
by: Hemmatyar, Mohammad Mahdi, et al.
Published: (2025)
by: Hemmatyar, Mohammad Mahdi, et al.
Published: (2025)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
REPAIR: Rank Correlation and Noisy Pair Half-replacing with Memory for Noisy Correspondence
by: Zheng, Ruochen, et al.
Published: (2024)
by: Zheng, Ruochen, et al.
Published: (2024)
SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere
by: Huang, Chao, et al.
Published: (2026)
by: Huang, Chao, et al.
Published: (2026)
Multi-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal Properties
by: Li, Wenqiao, et al.
Published: (2024)
by: Li, Wenqiao, et al.
Published: (2024)
RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection
by: Lee, Junhee, et al.
Published: (2025)
by: Lee, Junhee, et al.
Published: (2025)
Learn Suspected Anomalies from Event Prompts for Video Anomaly Detection
by: Tao, Chenchen, et al.
Published: (2024)
by: Tao, Chenchen, et al.
Published: (2024)
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
by: Li, Wenqiao, et al.
Published: (2025)
by: Li, Wenqiao, et al.
Published: (2025)
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
by: Deng, Haoyou, et al.
Published: (2026)
by: Deng, Haoyou, et al.
Published: (2026)
CLIP-guided Prototype Modulating for Few-shot Action Recognition
by: Wang, Xiang, et al.
Published: (2023)
by: Wang, Xiang, et al.
Published: (2023)
HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection
by: Cai, Zhaolin, et al.
Published: (2025)
by: Cai, Zhaolin, et al.
Published: (2025)
Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation
by: Wu, Dongyue, et al.
Published: (2024)
by: Wu, Dongyue, et al.
Published: (2024)
ProDisc-VAD: An Efficient System for Weakly-Supervised Anomaly Detection in Video Surveillance Applications
by: Zhu, Tao, et al.
Published: (2025)
by: Zhu, Tao, et al.
Published: (2025)
CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection
by: Lim, Hyeongmuk, et al.
Published: (2026)
by: Lim, Hyeongmuk, et al.
Published: (2026)
Adaptive Prototype Replay for Class Incremental Semantic Segmentation
by: Zhu, Guilin, et al.
Published: (2024)
by: Zhu, Guilin, et al.
Published: (2024)
SlowFastVAD: Video Anomaly Detection via Integrating Simple Detector and RAG-Enhanced Vision-Language Model
by: Ding, Zongcan, et al.
Published: (2025)
by: Ding, Zongcan, et al.
Published: (2025)
From Perfect to Noisy World Simulation: Customizable Embodied Multi-modal Perturbations for SLAM Robustness Benchmarking
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Similar Items
-
GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly Detection
by: Zhang, Huaxin, et al.
Published: (2024) -
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
by: Zhang, Huaxin, et al.
Published: (2024) -
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
by: Xu, Xiaohao, et al.
Published: (2024) -
Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment
by: Yin, Wenti, et al.
Published: (2025) -
HR-Pro: Point-supervised Temporal Action Localization via Hierarchical Reliability Propagation
by: Zhang, Huaxin, et al.
Published: (2023)