StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Junxi, Sun, Te, Zhu, Jiayi, Li, Junxian, Xu, Haowen, Wen, Zichen, Hu, Xuming, Li, Zhiyu, Zhang, Linfeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913057583661056
author Wang, Junxi
Sun, Te
Zhu, Jiayi
Li, Junxian
Xu, Haowen
Wen, Zichen
Hu, Xuming
Li, Zhiyu
Zhang, Linfeng
author_facet Wang, Junxi
Sun, Te
Zhu, Jiayi
Li, Junxian
Xu, Haowen
Wen, Zichen
Hu, Xuming
Li, Zhiyu
Zhang, Linfeng
contents Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memory overhead, leading to high costs in both storage and computation. To address this issue, we propose StreamMeCo, an efficient Stream Agent Memory Compression framework. Specifically, based on the connectivity of the memory graph, StreamMeCo introduces edge-free minmax sampling for the isolated nodes and an edge-aware weight pruning for connected nodes, evicting the redundant memory nodes while maintaining the accuracy. In addition, we introduce a time-decay memory retrieval mechanism to further eliminate the performance degradation caused by memory compression. Extensive experiments on three challenging benchmark datasets (M3-Bench-robot, M3-Bench-web and Video-MME-Long) demonstrate that under 70% memory graph compression, StreamMeCo achieves a 1.87* speedup in memory retrieval while delivering an average accuracy improvement of 1.0%. Our code is available at https://github.com/Celina-love-sweet/StreamMeCo.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09000
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
Wang, Junxi
Sun, Te
Zhu, Jiayi
Li, Junxian
Xu, Haowen
Wen, Zichen
Hu, Xuming
Li, Zhiyu
Zhang, Linfeng
Computer Vision and Pattern Recognition
Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memory overhead, leading to high costs in both storage and computation. To address this issue, we propose StreamMeCo, an efficient Stream Agent Memory Compression framework. Specifically, based on the connectivity of the memory graph, StreamMeCo introduces edge-free minmax sampling for the isolated nodes and an edge-aware weight pruning for connected nodes, evicting the redundant memory nodes while maintaining the accuracy. In addition, we introduce a time-decay memory retrieval mechanism to further eliminate the performance degradation caused by memory compression. Extensive experiments on three challenging benchmark datasets (M3-Bench-robot, M3-Bench-web and Video-MME-Long) demonstrate that under 70% memory graph compression, StreamMeCo achieves a 1.87* speedup in memory retrieval while delivering an average accuracy improvement of 1.0%. Our code is available at https://github.com/Celina-love-sweet/StreamMeCo.
title StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.09000