StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913057583661056 |
|---|---|
| author | Wang, Junxi Sun, Te Zhu, Jiayi Li, Junxian Xu, Haowen Wen, Zichen Hu, Xuming Li, Zhiyu Zhang, Linfeng |
| author_facet | Wang, Junxi Sun, Te Zhu, Jiayi Li, Junxian Xu, Haowen Wen, Zichen Hu, Xuming Li, Zhiyu Zhang, Linfeng |
| contents | Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memory overhead, leading to high costs in both storage and computation. To address this issue, we propose StreamMeCo, an efficient Stream Agent Memory Compression framework. Specifically, based on the connectivity of the memory graph, StreamMeCo introduces edge-free minmax sampling for the isolated nodes and an edge-aware weight pruning for connected nodes, evicting the redundant memory nodes while maintaining the accuracy. In addition, we introduce a time-decay memory retrieval mechanism to further eliminate the performance degradation caused by memory compression. Extensive experiments on three challenging benchmark datasets (M3-Bench-robot, M3-Bench-web and Video-MME-Long) demonstrate that under 70% memory graph compression, StreamMeCo achieves a 1.87* speedup in memory retrieval while delivering an average accuracy improvement of 1.0%. Our code is available at https://github.com/Celina-love-sweet/StreamMeCo. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_09000 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding Wang, Junxi Sun, Te Zhu, Jiayi Li, Junxian Xu, Haowen Wen, Zichen Hu, Xuming Li, Zhiyu Zhang, Linfeng Computer Vision and Pattern Recognition Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memory overhead, leading to high costs in both storage and computation. To address this issue, we propose StreamMeCo, an efficient Stream Agent Memory Compression framework. Specifically, based on the connectivity of the memory graph, StreamMeCo introduces edge-free minmax sampling for the isolated nodes and an edge-aware weight pruning for connected nodes, evicting the redundant memory nodes while maintaining the accuracy. In addition, we introduce a time-decay memory retrieval mechanism to further eliminate the performance degradation caused by memory compression. Extensive experiments on three challenging benchmark datasets (M3-Bench-robot, M3-Bench-web and Video-MME-Long) demonstrate that under 70% memory graph compression, StreamMeCo achieves a 1.87* speedup in memory retrieval while delivering an average accuracy improvement of 1.0%. Our code is available at https://github.com/Celina-love-sweet/StreamMeCo. |
| title | StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.09000 |