Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ye, Shengyuan, Ouyang, Bei, Qian, Tianyi, Zeng, Liekang, Yuan, Mu, Chu, Xiaowen, Hong, Weijie, Chen, Xu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914238402920448
author Ye, Shengyuan
Ouyang, Bei
Qian, Tianyi
Zeng, Liekang
Yuan, Mu
Chu, Xiaowen
Hong, Weijie
Chen, Xu
author_facet Ye, Shengyuan
Ouyang, Bei
Qian, Tianyi
Zeng, Liekang
Yuan, Mu
Chu, Xiaowen
Hong, Weijie
Chen, Xu
contents Vision-language models (VLMs) have demonstrated impressive multimodal comprehension capabilities and are being deployed in an increasing number of online video understanding applications. While recent efforts extensively explore advancing VLMs' reasoning power in these cases, deployment constraints are overlooked, leading to overwhelming system overhead in real-world deployments. To address that, we propose Venus, an on-device memory-and-retrieval system for efficient online video understanding. Venus proposes an edge-cloud disaggregated architecture that sinks memory construction and keyframe retrieval from cloud to edge, operating in two stages. In the ingestion stage, Venus continuously processes streaming edge videos via scene segmentation and clustering, where the selected keyframes are embedded with a multimodal embedding model to build a hierarchical memory for efficient storage and retrieval. In the querying stage, Venus indexes incoming queries from memory, and employs a threshold-based progressive sampling algorithm for keyframe selection that enhances diversity and adaptively balances system cost and reasoning accuracy. Our extensive evaluation shows that Venus achieves a 15x-131x speedup in total response latency compared to state-of-the-art methods, enabling real-time responses within seconds while maintaining comparable or even superior reasoning accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07344
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
Ye, Shengyuan
Ouyang, Bei
Qian, Tianyi
Zeng, Liekang
Yuan, Mu
Chu, Xiaowen
Hong, Weijie
Chen, Xu
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Vision-language models (VLMs) have demonstrated impressive multimodal comprehension capabilities and are being deployed in an increasing number of online video understanding applications. While recent efforts extensively explore advancing VLMs' reasoning power in these cases, deployment constraints are overlooked, leading to overwhelming system overhead in real-world deployments. To address that, we propose Venus, an on-device memory-and-retrieval system for efficient online video understanding. Venus proposes an edge-cloud disaggregated architecture that sinks memory construction and keyframe retrieval from cloud to edge, operating in two stages. In the ingestion stage, Venus continuously processes streaming edge videos via scene segmentation and clustering, where the selected keyframes are embedded with a multimodal embedding model to build a hierarchical memory for efficient storage and retrieval. In the querying stage, Venus indexes incoming queries from memory, and employs a threshold-based progressive sampling algorithm for keyframe selection that enhances diversity and adaptively balances system cost and reasoning accuracy. Our extensive evaluation shows that Venus achieves a 15x-131x speedup in total response latency compared to state-of-the-art methods, enabling real-time responses within seconds while maintaining comparable or even superior reasoning accuracy.
title Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2512.07344