StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911715580444672 |
|---|---|
| author | Xie, Ming Huang, Zizheng Tan, Xudong Wang, Chao Zeng, Xiangyu Wu, Wenxiao Chen, Tao Wang, Limin Fu, Yanwei |
| author_facet | Xie, Ming Huang, Zizheng Tan, Xudong Wang, Chao Zeng, Xiangyu Wu, Wenxiao Chen, Tao Wang, Limin Fu, Yanwei |
| contents | While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offline settings, limiting their applicability in streaming scenarios due to two fundamental flaws. First, they lack robust mechanisms to manage continuously growing audio-visual context over long horizons and cannot autonomously initiate responses at opportune moments. Second, existing benchmarks are predominantly confined to offline, single-turn question answering, failing to capture continuous, multi-turn streaming interactions. To bridge these gaps, we propose StreamOV, a novel Streaming Omni-Video understanding framework for efficient online audio-visual reasoning with bounded memory and proactive response triggering. Specifically, StreamOV introduces a multimodal evidence-guided long-short term memory that condenses historical audio-visual context into compact informative evidence under a fixed budget. It further employs a hidden-state-driven trigger to decide when to respond, avoiding explicit silence-token generation and external routers. We also curate SOVBench, the first comprehensive benchmark for online, multi-turn omni-modal evaluation. Extensive experiments show that StreamOV achieves state-of-the-art performance across diverse streaming and omni-video benchmarks, demonstrating its effectiveness for both online and offline video understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_25621 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Xie, Ming Huang, Zizheng Tan, Xudong Wang, Chao Zeng, Xiangyu Wu, Wenxiao Chen, Tao Wang, Limin Fu, Yanwei Computer Vision and Pattern Recognition While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offline settings, limiting their applicability in streaming scenarios due to two fundamental flaws. First, they lack robust mechanisms to manage continuously growing audio-visual context over long horizons and cannot autonomously initiate responses at opportune moments. Second, existing benchmarks are predominantly confined to offline, single-turn question answering, failing to capture continuous, multi-turn streaming interactions. To bridge these gaps, we propose StreamOV, a novel Streaming Omni-Video understanding framework for efficient online audio-visual reasoning with bounded memory and proactive response triggering. Specifically, StreamOV introduces a multimodal evidence-guided long-short term memory that condenses historical audio-visual context into compact informative evidence under a fixed budget. It further employs a hidden-state-driven trigger to decide when to respond, avoiding explicit silence-token generation and external routers. We also curate SOVBench, the first comprehensive benchmark for online, multi-turn omni-modal evaluation. Extensive experiments show that StreamOV achieves state-of-the-art performance across diverse streaming and omni-video benchmarks, demonstrating its effectiveness for both online and offline video understanding. |
| title | StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.25621 |