Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Huaxin, Xu, Xiaohao, Wang, Xiang, Zuo, Jialong, Huang, Xiaonan, Gao, Changxin, Zhang, Shanjun, Yu, Li, Sang, Nong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly Detection
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
HR-Pro: Point-supervised Temporal Action Localization via Hierarchical Reliability Propagation
by: Zhang, Huaxin, et al.
Published: (2023)
by: Zhang, Huaxin, et al.
Published: (2023)
Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment
by: Yin, Wenti, et al.
Published: (2025)
by: Yin, Wenti, et al.
Published: (2025)
UFineBench: Towards Text-based Person Retrieval with Ultra-fine Granularity
by: Zuo, Jialong, et al.
Published: (2023)
by: Zuo, Jialong, et al.
Published: (2023)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Cross-video Identity Correlating for Person Re-identification Pre-training
by: Zuo, Jialong, et al.
Published: (2024)
by: Zuo, Jialong, et al.
Published: (2024)
Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training Acceleration
by: Wu, Dongyue, et al.
Published: (2025)
by: Wu, Dongyue, et al.
Published: (2025)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
by: Erregue, Iñaki, et al.
Published: (2026)
by: Erregue, Iñaki, et al.
Published: (2026)
VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
by: Zhu, Liyun, et al.
Published: (2025)
by: Zhu, Liyun, et al.
Published: (2025)
PLIP: Language-Image Pre-training for Person Representation Learning
by: Zuo, Jialong, et al.
Published: (2023)
by: Zuo, Jialong, et al.
Published: (2023)
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
by: Pereira, João, et al.
Published: (2026)
by: Pereira, João, et al.
Published: (2026)
Object-Aware Video Matting with Cross-Frame Guidance
by: Zhang, Huayu, et al.
Published: (2025)
by: Zhang, Huayu, et al.
Published: (2025)
Spatial Cascaded Clustering and Weighted Memory for Unsupervised Person Re-identification
by: Hong, Jiahao, et al.
Published: (2024)
by: Hong, Jiahao, et al.
Published: (2024)
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
by: Huang, Yuzhi, et al.
Published: (2025)
by: Huang, Yuzhi, et al.
Published: (2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
REPAIR: Rank Correlation and Noisy Pair Half-replacing with Memory for Noisy Correspondence
by: Zheng, Ruochen, et al.
Published: (2024)
by: Zheng, Ruochen, et al.
Published: (2024)
DFIMat: Decoupled Flexible Interactive Matting in Multi-Person Scenarios
by: Jiao, Siyi, et al.
Published: (2024)
by: Jiao, Siyi, et al.
Published: (2024)
Replace Anyone in Videos
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation
by: Wu, Dongyue, et al.
Published: (2024)
by: Wu, Dongyue, et al.
Published: (2024)
CLIP-guided Prototype Modulating for Few-shot Action Recognition
by: Wang, Xiang, et al.
Published: (2023)
by: Wang, Xiang, et al.
Published: (2023)
SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time Segmentation
by: Xu, Zhengze, et al.
Published: (2023)
by: Xu, Zhengze, et al.
Published: (2023)
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
by: Li, Wenqiao, et al.
Published: (2025)
by: Li, Wenqiao, et al.
Published: (2025)
Segment Any Anomaly without Training via Hybrid Prompt Regularization
by: Cao, Yunkang, et al.
Published: (2023)
by: Cao, Yunkang, et al.
Published: (2023)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Photorealistic Phantom Roads in Real Scenes: Disentangling 3D Hallucinations from Physical Geometry
by: Nguyen, Hoang, et al.
Published: (2025)
by: Nguyen, Hoang, et al.
Published: (2025)
Latent Geometry Beyond Search: Amortizing Planning in World Models
by: Nguyen, Hoang, et al.
Published: (2026)
by: Nguyen, Hoang, et al.
Published: (2026)
A Survey on Visual Anomaly Detection: Challenge, Approach, and Prospect
by: Cao, Yunkang, et al.
Published: (2024)
by: Cao, Yunkang, et al.
Published: (2024)
Is Your LiDAR Placement Optimized for 3D Scene Understanding?
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
by: Deng, Haoyou, et al.
Published: (2026)
by: Deng, Haoyou, et al.
Published: (2026)
Taming Consistency Distillation for Accelerated Human Image Animation
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Lookup Table meets Local Laplacian Filter: Pyramid Reconstruction Network for Tone Mapping
by: Zhang, Feng, et al.
Published: (2023)
by: Zhang, Feng, et al.
Published: (2023)
Towards Understanding Camera Motions in Any Video
by: Lin, Zhiqiu, et al.
Published: (2025)
by: Lin, Zhiqiu, et al.
Published: (2025)
Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
by: Wu, Wenxiao, et al.
Published: (2025)
by: Wu, Wenxiao, et al.
Published: (2025)
Towards Ambiguity-Free Spatial Foundation Model: Rethinking and Decoupling Depth Ambiguity
by: Xu, Xiaohao, et al.
Published: (2025)
by: Xu, Xiaohao, et al.
Published: (2025)
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
by: Li, Jiahua, et al.
Published: (2025)
by: Li, Jiahua, et al.
Published: (2025)
Similar Items
-
Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
by: Zhang, Huaxin, et al.
Published: (2024) -
GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly Detection
by: Zhang, Huaxin, et al.
Published: (2024) -
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
by: Xu, Xiaohao, et al.
Published: (2024) -
HR-Pro: Point-supervised Temporal Action Localization via Hierarchical Reliability Propagation
by: Zhang, Huaxin, et al.
Published: (2023) -
Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment
by: Yin, Wenti, et al.
Published: (2025)