RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Sha, Prabhu, Yogesh, Ossowski, Timothy, Chen, Kaiping, Hu, Junjie
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908778378559488
author Luo, Sha
Prabhu, Yogesh
Ossowski, Timothy
Chen, Kaiping
Hu, Junjie
author_facet Luo, Sha
Prabhu, Yogesh
Ossowski, Timothy
Chen, Kaiping
Hu, Junjie
contents With the rapid growth of video centered social media, the ability to anticipate risky events from visual data is a promising direction for ensuring public safety and preventing real world accidents. Prior work has extensively studied supervised video risk assessment across domains such as driving, protests, and natural disasters. However, many existing datasets provide models with access to the full video sequence, including the accident itself, which substantially reduces the difficulty of the task. To better reflect real world conditions, we introduce a new video understanding benchmark RiskCueBench in which videos are carefully annotated to identify a risk signal clip, defined as the earliest moment that indicates a potential safety concern. Experimental results reveal a significant gap in current systems ability to interpret evolving situations and anticipate future risky events from early visual signals, highlighting important challenges for deploying video risk prediction models in practice.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03369
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
Luo, Sha
Prabhu, Yogesh
Ossowski, Timothy
Chen, Kaiping
Hu, Junjie
Computer Vision and Pattern Recognition
Computation and Language
With the rapid growth of video centered social media, the ability to anticipate risky events from visual data is a promising direction for ensuring public safety and preventing real world accidents. Prior work has extensively studied supervised video risk assessment across domains such as driving, protests, and natural disasters. However, many existing datasets provide models with access to the full video sequence, including the accident itself, which substantially reduces the difficulty of the task. To better reflect real world conditions, we introduce a new video understanding benchmark RiskCueBench in which videos are carefully annotated to identify a risk signal clip, defined as the earliest moment that indicates a potential safety concern. Experimental results reveal a significant gap in current systems ability to interpret evolving situations and anticipate future risky events from early visual signals, highlighting important challenges for deploying video risk prediction models in practice.
title RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2601.03369