CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Yulin, Chen, Yan, Chen, Wenyan, Park, JooYoung, Nitin, Shivaraman, Tao, Luo, Romero, Francisco, Ustiugov, Dmitrii |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
by: Chen, Wenyan, et al.
Published: (2026)
by: Chen, Wenyan, et al.
Published: (2026)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026)
by: Park, JooYoung, et al.
Published: (2026)
FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression
by: Mechels, Ben, et al.
Published: (2026)
by: Mechels, Ben, et al.
Published: (2026)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Building State Machine Replication Using Practical Network Synchrony
by: Wan, Yiliang, et al.
Published: (2025)
by: Wan, Yiliang, et al.
Published: (2025)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
by: Piao, XinYu, et al.
Published: (2021)
by: Piao, XinYu, et al.
Published: (2021)
VcLLM: Video Codecs are Secretly Tensor Codecs
by: Xu, Ceyu, et al.
Published: (2024)
by: Xu, Ceyu, et al.
Published: (2024)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
by: Mi, Liang, et al.
Published: (2026)
by: Mi, Liang, et al.
Published: (2026)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
by: Lai, Ruiqi, et al.
Published: (2025)
by: Lai, Ruiqi, et al.
Published: (2025)
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
Melding the Serverless Control Plane with the Conventional Cluster Manager for Speed and Resource Efficiency
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
by: Naman, Pranjal, et al.
Published: (2025)
by: Naman, Pranjal, et al.
Published: (2025)
WANSpec: Leveraging Global Compute Capacity for LLM Inference
by: Martin, Noah, et al.
Published: (2026)
by: Martin, Noah, et al.
Published: (2026)
DGNNFlow: A Streaming Dataflow Architecture for Real-Time Edge-based Dynamic GNN Inference in HL-LHC Trigger Systems
by: Maharaj, Davendra, et al.
Published: (2026)
by: Maharaj, Davendra, et al.
Published: (2026)
Cluster-based Network Time Synchronization for Resilience with Energy Efficiency
by: Shivaraman, Nitin, et al.
Published: (2024)
by: Shivaraman, Nitin, et al.
Published: (2024)
Collaborative Inference in DNN-based Satellite Systems with Dynamic Task Streams
by: Guan, Jinglong, et al.
Published: (2023)
by: Guan, Jinglong, et al.
Published: (2023)
Federated Inference for Heterogeneous LLM Communication and Collaboration
by: Chen, Zihan, et al.
Published: (2026)
by: Chen, Zihan, et al.
Published: (2026)
Byzantine Fault-Tolerant Min-Max Optimization
by: Liu, Shuo, et al.
Published: (2022)
by: Liu, Shuo, et al.
Published: (2022)
Leveraging Public Cloud Infrastructure for Real-time Connected Vehicle Speed Advisory at a Signalized Corridor
by: Deng, Hsien-Wen, et al.
Published: (2024)
by: Deng, Hsien-Wen, et al.
Published: (2024)
Byzantine Consensus in Directed Graphs with Message Authentication
by: Vaidya, Nitin H., et al.
Published: (2026)
by: Vaidya, Nitin H., et al.
Published: (2026)
Byzantine fault-tolerant distributed set intersection with redundancy
by: Liu, Shuo, et al.
Published: (2024)
by: Liu, Shuo, et al.
Published: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Optimizing Federated Learning in the Era of LLMs: Message Quantization and Streaming
by: Xu, Ziyue, et al.
Published: (2025)
by: Xu, Ziyue, et al.
Published: (2025)
RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
by: Ni, Xiaobing, et al.
Published: (2024)
by: Ni, Xiaobing, et al.
Published: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
Towards Exascale Computing for Astrophysical Simulation Leveraging the Leonardo EuroHPC System
by: Shukla, Nitin, et al.
Published: (2025)
by: Shukla, Nitin, et al.
Published: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
by: Wang, Weiye, et al.
Published: (2026)
by: Wang, Weiye, et al.
Published: (2026)
Big Data-Driven Fraud Detection Using Machine Learning and Real-Time Stream Processing
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
Asynchronous Checkpoint for Eventually Consistent Databases
by: Ravishankar, Raaghav, et al.
Published: (2025)
by: Ravishankar, Raaghav, et al.
Published: (2025)
Approximate Byzantine Fault-Tolerance in Distributed Optimization
by: Liu, Shuo, et al.
Published: (2021)
by: Liu, Shuo, et al.
Published: (2021)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
by: Chen, David, et al.
Published: (2026)
by: Chen, David, et al.
Published: (2026)
KV Cache Compression for Inference Efficiency in LLMs: A Review
by: Liu, Yanyu, et al.
Published: (2025)
by: Liu, Yanyu, et al.
Published: (2025)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
Stream-K Optimization and Exploration
by: Rackley, Nick, et al.
Published: (2024)
by: Rackley, Nick, et al.
Published: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
by: Chen, Xing, et al.
Published: (2025)
by: Chen, Xing, et al.
Published: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
Similar Items
-
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
by: Chen, Wenyan, et al.
Published: (2026) -
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026) -
FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression
by: Mechels, Ben, et al.
Published: (2026) -
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026) -
Building State Machine Replication Using Practical Network Synchrony
by: Wan, Yiliang, et al.
Published: (2025)