StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Ming, Huang, Zizheng, Tan, Xudong, Wang, Chao, Zeng, Xiangyu, Wu, Wenxiao, Chen, Tao, Wang, Limin, Fu, Yanwei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911715580444672
author Xie, Ming
Huang, Zizheng
Tan, Xudong
Wang, Chao
Zeng, Xiangyu
Wu, Wenxiao
Chen, Tao
Wang, Limin
Fu, Yanwei
author_facet Xie, Ming
Huang, Zizheng
Tan, Xudong
Wang, Chao
Zeng, Xiangyu
Wu, Wenxiao
Chen, Tao
Wang, Limin
Fu, Yanwei
contents While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offline settings, limiting their applicability in streaming scenarios due to two fundamental flaws. First, they lack robust mechanisms to manage continuously growing audio-visual context over long horizons and cannot autonomously initiate responses at opportune moments. Second, existing benchmarks are predominantly confined to offline, single-turn question answering, failing to capture continuous, multi-turn streaming interactions. To bridge these gaps, we propose StreamOV, a novel Streaming Omni-Video understanding framework for efficient online audio-visual reasoning with bounded memory and proactive response triggering. Specifically, StreamOV introduces a multimodal evidence-guided long-short term memory that condenses historical audio-visual context into compact informative evidence under a fixed budget. It further employs a hidden-state-driven trigger to decide when to respond, avoiding explicit silence-token generation and external routers. We also curate SOVBench, the first comprehensive benchmark for online, multi-turn omni-modal evaluation. Extensive experiments show that StreamOV achieves state-of-the-art performance across diverse streaming and omni-video benchmarks, demonstrating its effectiveness for both online and offline video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25621
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Xie, Ming
Huang, Zizheng
Tan, Xudong
Wang, Chao
Zeng, Xiangyu
Wu, Wenxiao
Chen, Tao
Wang, Limin
Fu, Yanwei
Computer Vision and Pattern Recognition
While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offline settings, limiting their applicability in streaming scenarios due to two fundamental flaws. First, they lack robust mechanisms to manage continuously growing audio-visual context over long horizons and cannot autonomously initiate responses at opportune moments. Second, existing benchmarks are predominantly confined to offline, single-turn question answering, failing to capture continuous, multi-turn streaming interactions. To bridge these gaps, we propose StreamOV, a novel Streaming Omni-Video understanding framework for efficient online audio-visual reasoning with bounded memory and proactive response triggering. Specifically, StreamOV introduces a multimodal evidence-guided long-short term memory that condenses historical audio-visual context into compact informative evidence under a fixed budget. It further employs a hidden-state-driven trigger to decide when to respond, avoiding explicit silence-token generation and external routers. We also curate SOVBench, the first comprehensive benchmark for online, multi-turn omni-modal evaluation. Extensive experiments show that StreamOV achieves state-of-the-art performance across diverse streaming and omni-video benchmarks, demonstrating its effectiveness for both online and offline video understanding.
title StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.25621