MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Shiju, Hu, Junhao, Huang, Rongxiao, Zheng, Jiaqi, Chen, Guihai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
You Need an Encoder for Native Position-Independent Caching
von: Zhao, Shiju, et al.
Veröffentlicht: (2026)
von: Zhao, Shiju, et al.
Veröffentlicht: (2026)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
Reinfier and Reintrainer: Verification and Interpretation-Driven Safe Deep Reinforcement Learning Frameworks
von: Yang, Zixuan, et al.
Veröffentlicht: (2024)
von: Yang, Zixuan, et al.
Veröffentlicht: (2024)
IC-Cache: Efficient Large Language Model Serving via In-context Caching
von: Yu, Yifan, et al.
Veröffentlicht: (2025)
von: Yu, Yifan, et al.
Veröffentlicht: (2025)
Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management
von: Zheng, Haoyu, et al.
Veröffentlicht: (2026)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2026)
MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
von: Gao, Shihong, et al.
Veröffentlicht: (2025)
von: Gao, Shihong, et al.
Veröffentlicht: (2025)
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
von: Kim, Jungwoo, et al.
Veröffentlicht: (2025)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
Decision Focused Causal Learning for Direct Counterfactual Marketing Optimization
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
von: Zeng, Hui, et al.
Veröffentlicht: (2025)
von: Zeng, Hui, et al.
Veröffentlicht: (2025)
Beyond Anti-Forgetting: Multimodal Continual Instruction Tuning with Positive Forward Transfer
von: Zheng, Junhao, et al.
Veröffentlicht: (2024)
von: Zheng, Junhao, et al.
Veröffentlicht: (2024)
DGHMesh: A Large-scale Dual-radar mmWave Dataset and Generalization-focused Benchmark for Human Mesh Reconstruction
von: Guo, Rongxiao, et al.
Veröffentlicht: (2026)
von: Guo, Rongxiao, et al.
Veröffentlicht: (2026)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
von: Yao, Jiayi, et al.
Veröffentlicht: (2024)
von: Yao, Jiayi, et al.
Veröffentlicht: (2024)
Revisiting Service Level Objectives and System Level Metrics in Large Language Model Serving
von: Wang, Zhibin, et al.
Veröffentlicht: (2024)
von: Wang, Zhibin, et al.
Veröffentlicht: (2024)
Continuous Semantic Caching for Low-Cost LLM Serving
von: Atalar, Baran, et al.
Veröffentlicht: (2026)
von: Atalar, Baran, et al.
Veröffentlicht: (2026)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
von: Gao, Bin, et al.
Veröffentlicht: (2024)
von: Gao, Bin, et al.
Veröffentlicht: (2024)
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
Bi-Level Decision-Focused Causal Learning for Large-Scale Marketing Optimization: Bridging Observational and Experimental Data
von: Zhang, Shuli, et al.
Veröffentlicht: (2025)
von: Zhang, Shuli, et al.
Veröffentlicht: (2025)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
von: Zeng, Pai, et al.
Veröffentlicht: (2024)
von: Zeng, Pai, et al.
Veröffentlicht: (2024)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
von: Zhao, Yilong, et al.
Veröffentlicht: (2023)
von: Zhao, Yilong, et al.
Veröffentlicht: (2023)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
von: Shirokikh, Mikhail, et al.
Veröffentlicht: (2026)
von: Shirokikh, Mikhail, et al.
Veröffentlicht: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
Scaling Graph Chain-of-Thought Reasoning: A Multi-Agent Framework with Efficient LLM Serving
von: Huan, Chengying, et al.
Veröffentlicht: (2025)
von: Huan, Chengying, et al.
Veröffentlicht: (2025)
ChemMLLM: Chemical Multimodal Large Language Model
von: Tan, Qian, et al.
Veröffentlicht: (2025)
von: Tan, Qian, et al.
Veröffentlicht: (2025)
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
iServe: An Intent-based Serving System for LLMs
von: Liakopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Liakopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
von: Jia, Jinda, et al.
Veröffentlicht: (2026)
von: Jia, Jinda, et al.
Veröffentlicht: (2026)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
von: Yamamoto, Yuji, et al.
Veröffentlicht: (2026)
von: Yamamoto, Yuji, et al.
Veröffentlicht: (2026)
Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning
von: Liu, Jiajin, et al.
Veröffentlicht: (2025)
von: Liu, Jiajin, et al.
Veröffentlicht: (2025)
Joint Encoding of KV-Cache Blocks for Scalable LLM Serving
von: Kampeas, Joseph, et al.
Veröffentlicht: (2026)
von: Kampeas, Joseph, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
You Need an Encoder for Native Position-Independent Caching
von: Zhao, Shiju, et al.
Veröffentlicht: (2026) -
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024) -
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
von: Wang, Qian, et al.
Veröffentlicht: (2025) -
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
von: Ma, Bole, et al.
Veröffentlicht: (2026) -
Reinfier and Reintrainer: Verification and Interpretation-Driven Safe Deep Reinforcement Learning Frameworks
von: Yang, Zixuan, et al.
Veröffentlicht: (2024)