Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Vatsal, Suri, Saksham, Gwilliam, Matthew, Kumar, Pulkit, Shrivastava, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025)
by: Aggarwal, Anirud, et al.
Published: (2025)
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025)
by: Agarwal, Vatsal, et al.
Published: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024)
by: Wang, Hanyu, et al.
Published: (2024)
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026)
by: Walmer, Matthew, et al.
Published: (2026)
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)
by: Suri, Saksham, et al.
Published: (2024)
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026)
by: Padmanabhan, Namitha, et al.
Published: (2026)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
by: Zhang, Haowei, et al.
Published: (2026)
by: Zhang, Haowei, et al.
Published: (2026)
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024)
by: Padmanabhan, Namitha, et al.
Published: (2024)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
by: Jayatilaka, Gihan, et al.
Published: (2025)
by: Jayatilaka, Gihan, et al.
Published: (2025)
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026)
by: Das, Gautom, et al.
Published: (2026)
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)
by: Kumar, Pulkit, et al.
Published: (2025)
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
Do text-free diffusion models learn discriminative visual representations?
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
by: Maiya, Shishira R, et al.
Published: (2024)
by: Maiya, Shishira R, et al.
Published: (2024)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2024)
by: Kumar, Pulkit, et al.
Published: (2024)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
by: Patel, Shrenik, et al.
Published: (2025)
by: Patel, Shrenik, et al.
Published: (2025)
Characterizing Motion Encoding in Video Diffusion Timesteps
by: Baherwani, Vatsal, et al.
Published: (2025)
by: Baherwani, Vatsal, et al.
Published: (2025)
NeRF-Aug: Data Augmentation for Robotics with Neural Radiance Fields
by: Zhu, Eric, et al.
Published: (2024)
by: Zhu, Eric, et al.
Published: (2024)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
by: Gwilliam, Matthew, et al.
Published: (2023)
by: Gwilliam, Matthew, et al.
Published: (2023)
Accelerate High-Quality Diffusion Models with Inner Loop Feedback
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
by: Chen, Yilong, et al.
Published: (2025)
by: Chen, Yilong, et al.
Published: (2025)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
by: Ning, Zhenyu, et al.
Published: (2025)
by: Ning, Zhenyu, et al.
Published: (2025)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
by: He, Bo, et al.
Published: (2024)
by: He, Bo, et al.
Published: (2024)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation
by: Swaminathan, Archana, et al.
Published: (2024)
by: Swaminathan, Archana, et al.
Published: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
by: Chatterjee, Dibyadip, et al.
Published: (2025)
by: Chatterjee, Dibyadip, et al.
Published: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
by: Zeng, Xiangyu, et al.
Published: (2025)
by: Zeng, Xiangyu, et al.
Published: (2025)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
by: Xie, Ming, et al.
Published: (2026)
by: Xie, Ming, et al.
Published: (2026)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
by: Wang, Junxi, et al.
Published: (2026)
by: Wang, Junxi, et al.
Published: (2026)
Similar Items
-
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025) -
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025) -
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025) -
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024) -
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026)