LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ning, Zhenyu, Liu, Guangda, Jin, Qihao, Li, Chengwei, Ding, Wenchao, Guo, Minyi, Zhao, Jieru |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
von: Ning, Zhenyu, et al.
Veröffentlicht: (2024)
von: Ning, Zhenyu, et al.
Veröffentlicht: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Support
von: Zhang, Chenqi, et al.
Veröffentlicht: (2025)
von: Zhang, Chenqi, et al.
Veröffentlicht: (2025)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
SparseTem: Boosting the Efficiency of CNN-Based Video Encoders by Exploiting Temporal Continuity
von: Wang, Kunyun, et al.
Veröffentlicht: (2024)
von: Wang, Kunyun, et al.
Veröffentlicht: (2024)
LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
von: Yang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Yang, Zhenyu, et al.
Veröffentlicht: (2025)
V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache Retrieval
von: Kim, Donghyuk, et al.
Veröffentlicht: (2025)
von: Kim, Donghyuk, et al.
Veröffentlicht: (2025)
Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
von: Wang, Kunyun, et al.
Veröffentlicht: (2025)
von: Wang, Kunyun, et al.
Veröffentlicht: (2025)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic Termination
von: Feng, Yu, et al.
Veröffentlicht: (2025)
von: Feng, Yu, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture
von: Feng, Yu, et al.
Veröffentlicht: (2024)
von: Feng, Yu, et al.
Veröffentlicht: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
Hierarchical Source-to-Post-Route QoR Prediction in High-Level Synthesis with GNNs
von: Gao, Mingzhe, et al.
Veröffentlicht: (2024)
von: Gao, Mingzhe, et al.
Veröffentlicht: (2024)
AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMs
von: Gao, Mingzhe, et al.
Veröffentlicht: (2024)
von: Gao, Mingzhe, et al.
Veröffentlicht: (2024)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
von: Zhang, Qiuyang, et al.
Veröffentlicht: (2026)
von: Zhang, Qiuyang, et al.
Veröffentlicht: (2026)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
von: Mi, Liang, et al.
Veröffentlicht: (2026)
von: Mi, Liang, et al.
Veröffentlicht: (2026)
Online Misinformation Detection in Live Streaming Videos
von: Cao, Rui
Veröffentlicht: (2025)
von: Cao, Rui
Veröffentlicht: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025) -
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
von: Ning, Zhenyu, et al.
Veröffentlicht: (2024) -
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024) -
STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Support
von: Zhang, Chenqi, et al.
Veröffentlicht: (2025) -
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)