Strata: Hierarchical Context Caching for Long Context Language Model Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Zhiqiang, Xu, Ziyi, Zhao, Mark, An, Yuwei, Mailthody, Vikram Sharma, Mahlke, Scott, Garland, Michael, Kozyrakis, Christos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FailSafe: High-performance Resilient Serving
by: Xu, Ziyi, et al.
Published: (2025)
by: Xu, Ziyi, et al.
Published: (2025)
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
by: Park, Jeongmin Brian, et al.
Published: (2024)
by: Park, Jeongmin Brian, et al.
Published: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
by: Zhao, Mark, et al.
Published: (2024)
by: Zhao, Mark, et al.
Published: (2024)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
by: Yoshimura, Takeshi, et al.
Published: (2026)
by: Yoshimura, Takeshi, et al.
Published: (2026)
Regulating Branch Parallelism in LLM Serving
by: Gandhi, Swapnil, et al.
Published: (2026)
by: Gandhi, Swapnil, et al.
Published: (2026)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
by: Wu, Bingyang, et al.
Published: (2024)
by: Wu, Bingyang, et al.
Published: (2024)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
by: Wu, Bingyang, et al.
Published: (2025)
by: Wu, Bingyang, et al.
Published: (2025)
Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
by: Xie, Zhiqiang, et al.
Published: (2024)
by: Xie, Zhiqiang, et al.
Published: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
by: He, Yiyuan, et al.
Published: (2025)
by: He, Yiyuan, et al.
Published: (2025)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
by: Liu, Kaiwei, et al.
Published: (2025)
by: Liu, Kaiwei, et al.
Published: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
by: Nian, Sean, et al.
Published: (2026)
by: Nian, Sean, et al.
Published: (2026)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
by: Skiadopoulos, Athinagoras, et al.
Published: (2025)
by: Skiadopoulos, Athinagoras, et al.
Published: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
by: Zhang, Huawei, et al.
Published: (2025)
by: Zhang, Huawei, et al.
Published: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
by: Qianli, Liu, et al.
Published: (2025)
by: Qianli, Liu, et al.
Published: (2025)
Efficient Long Context Fine-tuning with Chunk Flow
by: Yuan, Xiulong, et al.
Published: (2025)
by: Yuan, Xiulong, et al.
Published: (2025)
FedCache: A Knowledge Cache-driven Federated Learning Architecture for Personalized Edge Intelligence
by: Wu, Zhiyuan, et al.
Published: (2023)
by: Wu, Zhiyuan, et al.
Published: (2023)
InstCache: A Predictive Cache for LLM Serving
by: Zou, Longwei, et al.
Published: (2024)
by: Zou, Longwei, et al.
Published: (2024)
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
by: Xie, Zhiqiang, et al.
Published: (2024)
by: Xie, Zhiqiang, et al.
Published: (2024)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
by: Tian, Yuyang, et al.
Published: (2025)
by: Tian, Yuyang, et al.
Published: (2025)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
by: Yuan, Ying, et al.
Published: (2026)
by: Yuan, Ying, et al.
Published: (2026)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026)
by: Wang, Wenfeng, et al.
Published: (2026)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
by: Zhu, Kan, et al.
Published: (2024)
by: Zhu, Kan, et al.
Published: (2024)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
by: Zhu, Jianian, et al.
Published: (2025)
by: Zhu, Jianian, et al.
Published: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
by: Gu, Diandian, et al.
Published: (2024)
by: Gu, Diandian, et al.
Published: (2024)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
by: Yoon, Dongha, et al.
Published: (2025)
by: Yoon, Dongha, et al.
Published: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026)
by: Bian, Zhuohang, et al.
Published: (2026)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
by: Poptani, Akash, et al.
Published: (2025)
by: Poptani, Akash, et al.
Published: (2025)
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
by: Yi, Mengjun, et al.
Published: (2025)
by: Yi, Mengjun, et al.
Published: (2025)
SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training
by: Wu, Kun, et al.
Published: (2024)
by: Wu, Kun, et al.
Published: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
by: Zhang, Yuning, et al.
Published: (2025)
by: Zhang, Yuning, et al.
Published: (2025)
FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving
by: Shen, Ao, et al.
Published: (2024)
by: Shen, Ao, et al.
Published: (2024)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
by: Liang, Yan, et al.
Published: (2026)
by: Liang, Yan, et al.
Published: (2026)
Similar Items
-
FailSafe: High-performance Resilient Serving
by: Xu, Ziyi, et al.
Published: (2025) -
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
by: Park, Jeongmin Brian, et al.
Published: (2024) -
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024) -
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024) -
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)