Don't Pause! Every prediction matters in a streaming video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chatterjee, Dibyadip, Pang, Zhanzhong, Sener, Fadime, Song, Yale, Yao, Angela
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910169321963520
author Chatterjee, Dibyadip
Pang, Zhanzhong
Sener, Fadime
Song, Yale
Yao, Angela
author_facet Chatterjee, Dibyadip
Pang, Zhanzhong
Sener, Fadime
Song, Yale
Yao, Angela
contents Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause the video at fixed timestamps, pose questions about current or past events, and score models only at those moments. This protocol leaves streaming predictions untested. To close this gap, we introduce SPOT-Bench, featuring multi-turn proactive queries that evaluate general streaming perception and assistive capabilities required by an always-on, real-time assistant. SPOT-Bench comes with Timeliness-F1, a consolidated metric that measures streaming predictions by their temporal precision and balanced coverage across the entire video. Our benchmark reveals: (i) offline models detect events reliably but spam predictions unprompted; (ii) post-training for silence reduces spamming but induces unresponsiveness; (iii) half of the streaming video expects no response, which we term dead-time - compute spent here does not affect response latency. These findings motivate AsynKV, a training-free streaming adaptation of offline models, that retains their event perception while improving their streaming behavior. AsynKV features a long-short term memory, utilized efficiently by scaling compute during dead-time. It serves as a strong baseline on SPOT-Bench, outperforming existing streaming models, and achieves state-of-the-art on retrospective benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24317
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Don't Pause! Every prediction matters in a streaming video
Chatterjee, Dibyadip
Pang, Zhanzhong
Sener, Fadime
Song, Yale
Yao, Angela
Computer Vision and Pattern Recognition
Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause the video at fixed timestamps, pose questions about current or past events, and score models only at those moments. This protocol leaves streaming predictions untested. To close this gap, we introduce SPOT-Bench, featuring multi-turn proactive queries that evaluate general streaming perception and assistive capabilities required by an always-on, real-time assistant. SPOT-Bench comes with Timeliness-F1, a consolidated metric that measures streaming predictions by their temporal precision and balanced coverage across the entire video. Our benchmark reveals: (i) offline models detect events reliably but spam predictions unprompted; (ii) post-training for silence reduces spamming but induces unresponsiveness; (iii) half of the streaming video expects no response, which we term dead-time - compute spent here does not affect response latency. These findings motivate AsynKV, a training-free streaming adaptation of offline models, that retains their event perception while improving their streaming behavior. AsynKV features a long-short term memory, utilized efficiently by scaling compute during dead-time. It serves as a strong baseline on SPOT-Bench, outperforming existing streaming models, and achieves state-of-the-art on retrospective benchmarks.
title Don't Pause! Every prediction matters in a streaming video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.24317