CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Yating, Cao, Congqi, Wang, Zhaoying, Meng, Weihua, Li, Jie, Li, Yuxin, Wei, Zihao, Shen, Zhongpei, Zhang, Jiajun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909882066665472
author Yu, Yating
Cao, Congqi
Wang, Zhaoying
Meng, Weihua
Li, Jie
Li, Yuxin
Wei, Zihao
Shen, Zhongpei
Zhang, Jiajun
author_facet Yu, Yating
Cao, Congqi
Wang, Zhaoying
Meng, Weihua
Li, Jie
Li, Yuxin
Wei, Zihao
Shen, Zhongpei
Zhang, Jiajun
contents How far are deep models from real-world video anomaly understanding (VAU)? Current works typically emphasize on detecting unexpected occurrences deviated from normal patterns or comprehending anomalous events with interpretable descriptions. However, they exhibit only a superficial comprehension of real-world anomalies, with limited breadth in complex principles and subtle context that distinguish the anomalies from normalities, e.g., climbing cliffs with safety gear vs. without it. To this end, we introduce CueBench, the first of its kind Benchmark, devoted to Context-aware video anomalies within a Unified Evaluation framework. We comprehensively establish an event-centric hierarchical taxonomy that anchors two core event types: 14 conditional and 18 absolute anomaly events, defined by their refined semantics from diverse contexts across 174 scenes and 198 attributes. Based on this, we propose to unify and benchmark context-aware VAU with various challenging tasks across recognition, temporal grounding, detection, and anticipation. This also serves as a rigorous and fair probing evaluation suite for generative-discriminative as well as generalized-specialized vision-language models (VLMs). To address the challenges underlying CueBench, we further develop Cue-R1 based on R1-style reinforcement fine-tuning with verifiable, task-aligned, and hierarchy-refined rewards in a unified generative manner. Extensive results on CueBench reveal that, existing VLMs are still far from satisfactory real-world anomaly understanding, while our Cue-R1 surpasses these state-of-the-art approaches by over 24% on average.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
Yu, Yating
Cao, Congqi
Wang, Zhaoying
Meng, Weihua
Li, Jie
Li, Yuxin
Wei, Zihao
Shen, Zhongpei
Zhang, Jiajun
Computer Vision and Pattern Recognition
How far are deep models from real-world video anomaly understanding (VAU)? Current works typically emphasize on detecting unexpected occurrences deviated from normal patterns or comprehending anomalous events with interpretable descriptions. However, they exhibit only a superficial comprehension of real-world anomalies, with limited breadth in complex principles and subtle context that distinguish the anomalies from normalities, e.g., climbing cliffs with safety gear vs. without it. To this end, we introduce CueBench, the first of its kind Benchmark, devoted to Context-aware video anomalies within a Unified Evaluation framework. We comprehensively establish an event-centric hierarchical taxonomy that anchors two core event types: 14 conditional and 18 absolute anomaly events, defined by their refined semantics from diverse contexts across 174 scenes and 198 attributes. Based on this, we propose to unify and benchmark context-aware VAU with various challenging tasks across recognition, temporal grounding, detection, and anticipation. This also serves as a rigorous and fair probing evaluation suite for generative-discriminative as well as generalized-specialized vision-language models (VLMs). To address the challenges underlying CueBench, we further develop Cue-R1 based on R1-style reinforcement fine-tuning with verifiable, task-aligned, and hierarchy-refined rewards in a unified generative manner. Extensive results on CueBench reveal that, existing VLMs are still far from satisfactory real-world anomaly understanding, while our Cue-R1 surpasses these state-of-the-art approaches by over 24% on average.
title CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.00613