StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yanlai, Zhao, Zhuokai, Shukla, Satya Narayan, Singh, Aashu, Mishra, Shlok Kumar, Zhang, Lizhu, Ren, Mengye
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909747803848704
author Yang, Yanlai
Zhao, Zhuokai
Shukla, Satya Narayan
Singh, Aashu
Mishra, Shlok Kumar
Zhang, Lizhu
Ren, Mengye
author_facet Yang, Yanlai
Zhao, Zhuokai
Shukla, Satya Narayan
Singh, Aashu
Mishra, Shlok Kumar
Zhang, Lizhu
Ren, Mengye
contents Multimodal large language models (MLLMs) have made significant progress in visual-language reasoning, but their ability to efficiently handle long videos remains limited. Despite recent advances in long-context MLLMs, storing and attending to the key-value (KV) cache for long visual contexts incurs substantial memory and computational overhead. Existing visual compression methods require either encoding the entire visual context before compression or having access to the questions in advance, which is impractical for long video understanding and multi-turn conversational settings. In this work, we propose StreamMem, a query-agnostic KV cache memory mechanism for streaming video understanding. Specifically, StreamMem encodes new video frames in a streaming manner, compressing the KV cache using attention scores between visual tokens and generic query tokens, while maintaining a fixed-size KV memory to enable efficient question answering (QA) in memory-constrained, long-video scenarios. Evaluation on three long video understanding and two streaming video question answering benchmarks shows that StreamMem achieves state-of-the-art performance in query-agnostic KV cache compression and is competitive with query-aware compression approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
Yang, Yanlai
Zhao, Zhuokai
Shukla, Satya Narayan
Singh, Aashu
Mishra, Shlok Kumar
Zhang, Lizhu
Ren, Mengye
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal large language models (MLLMs) have made significant progress in visual-language reasoning, but their ability to efficiently handle long videos remains limited. Despite recent advances in long-context MLLMs, storing and attending to the key-value (KV) cache for long visual contexts incurs substantial memory and computational overhead. Existing visual compression methods require either encoding the entire visual context before compression or having access to the questions in advance, which is impractical for long video understanding and multi-turn conversational settings. In this work, we propose StreamMem, a query-agnostic KV cache memory mechanism for streaming video understanding. Specifically, StreamMem encodes new video frames in a streaming manner, compressing the KV cache using attention scores between visual tokens and generic query tokens, while maintaining a fixed-size KV memory to enable efficient question answering (QA) in memory-constrained, long-video scenarios. Evaluation on three long video understanding and two streaming video question answering benchmarks shows that StreamMem achieves state-of-the-art performance in query-agnostic KV cache compression and is competitive with query-aware compression approaches.
title StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.15717