LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
Fuente:
arXiv
Guardado en:
| Autores principales: | Ning, Zhenyu, Liu, Guangda, Jin, Qihao, Li, Chengwei, Ding, Wenchao, Guo, Minyi, Zhao, Jieru |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
por: Ning, Zhenyu, et al.
Publicado: (2024)
por: Ning, Zhenyu, et al.
Publicado: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
por: Liu, Guangda, et al.
Publicado: (2024)
por: Liu, Guangda, et al.
Publicado: (2024)
STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Support
por: Zhang, Chenqi, et al.
Publicado: (2025)
por: Zhang, Chenqi, et al.
Publicado: (2025)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
por: Pang, Zhanzhong, et al.
Publicado: (2026)
por: Pang, Zhanzhong, et al.
Publicado: (2026)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
por: Zhang, Haowei, et al.
Publicado: (2026)
por: Zhang, Haowei, et al.
Publicado: (2026)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
por: Yang, Yanlai, et al.
Publicado: (2025)
por: Yang, Yanlai, et al.
Publicado: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
por: Chen, Yilong, et al.
Publicado: (2025)
por: Chen, Yilong, et al.
Publicado: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
por: Di, Shangzhe, et al.
Publicado: (2025)
por: Di, Shangzhe, et al.
Publicado: (2025)
SparseTem: Boosting the Efficiency of CNN-Based Video Encoders by Exploiting Temporal Continuity
por: Wang, Kunyun, et al.
Publicado: (2024)
por: Wang, Kunyun, et al.
Publicado: (2024)
LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
por: Yang, Zhenyu, et al.
Publicado: (2025)
por: Yang, Zhenyu, et al.
Publicado: (2025)
V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache Retrieval
por: Kim, Donghyuk, et al.
Publicado: (2025)
por: Kim, Donghyuk, et al.
Publicado: (2025)
Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
por: Wang, Kunyun, et al.
Publicado: (2025)
por: Wang, Kunyun, et al.
Publicado: (2025)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
por: Ye, Shengyuan, et al.
Publicado: (2025)
por: Ye, Shengyuan, et al.
Publicado: (2025)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
por: Zhang, Hang, et al.
Publicado: (2025)
por: Zhang, Hang, et al.
Publicado: (2025)
StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic Termination
por: Feng, Yu, et al.
Publicado: (2025)
por: Feng, Yu, et al.
Publicado: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
por: Xiao, Junbin, et al.
Publicado: (2026)
por: Xiao, Junbin, et al.
Publicado: (2026)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
por: Agarwal, Vatsal, et al.
Publicado: (2026)
por: Agarwal, Vatsal, et al.
Publicado: (2026)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
por: Wang, Tianze, et al.
Publicado: (2025)
por: Wang, Tianze, et al.
Publicado: (2025)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
por: Ji, Yicheng, et al.
Publicado: (2026)
por: Ji, Yicheng, et al.
Publicado: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
por: Wu, Wenbo, et al.
Publicado: (2025)
por: Wu, Wenbo, et al.
Publicado: (2025)
Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture
por: Feng, Yu, et al.
Publicado: (2024)
por: Feng, Yu, et al.
Publicado: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
por: Xu, Ruyi, et al.
Publicado: (2025)
por: Xu, Ruyi, et al.
Publicado: (2025)
Hierarchical Source-to-Post-Route QoR Prediction in High-Level Synthesis with GNNs
por: Gao, Mingzhe, et al.
Publicado: (2024)
por: Gao, Mingzhe, et al.
Publicado: (2024)
AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMs
por: Gao, Mingzhe, et al.
Publicado: (2024)
por: Gao, Mingzhe, et al.
Publicado: (2024)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
por: Zhang, Weichuang, et al.
Publicado: (2024)
por: Zhang, Weichuang, et al.
Publicado: (2024)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
por: Zhang, Qiuyang, et al.
Publicado: (2026)
por: Zhang, Qiuyang, et al.
Publicado: (2026)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
por: Patel, Shrenik, et al.
Publicado: (2025)
por: Patel, Shrenik, et al.
Publicado: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026)
por: Ji, Shiyu, et al.
Publicado: (2026)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
por: Gu, Yifeng, et al.
Publicado: (2025)
por: Gu, Yifeng, et al.
Publicado: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
por: Zeng, Xiangyu, et al.
Publicado: (2025)
por: Zeng, Xiangyu, et al.
Publicado: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
por: Mi, Liang, et al.
Publicado: (2026)
por: Mi, Liang, et al.
Publicado: (2026)
Online Misinformation Detection in Live Streaming Videos
por: Cao, Rui
Publicado: (2025)
por: Cao, Rui
Publicado: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
por: Liu, Yuhan, et al.
Publicado: (2023)
por: Liu, Yuhan, et al.
Publicado: (2023)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
por: Wang, Shihao, et al.
Publicado: (2026)
por: Wang, Shihao, et al.
Publicado: (2026)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
por: Jiang, Zhonghua, et al.
Publicado: (2025)
por: Jiang, Zhonghua, et al.
Publicado: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
Ejemplares similares
-
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025) -
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
por: Ning, Zhenyu, et al.
Publicado: (2024) -
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
por: Liu, Guangda, et al.
Publicado: (2024) -
STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Support
por: Zhang, Chenqi, et al.
Publicado: (2025) -
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
por: Pang, Zhanzhong, et al.
Publicado: (2026)