Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Ruoyu, He, Weiran, Wang, Yaoyu, Li, Zheming, Xu, Xinran, Wu, Yongwei, Zheng, Weimin, Zhang, Mingxing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
by: Qin, Ruoyu, et al.
Published: (2024)
by: Qin, Ruoyu, et al.
Published: (2024)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
by: Qin, Ruoyu, et al.
Published: (2025)
by: Qin, Ruoyu, et al.
Published: (2025)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)
by: Chen, Tiancheng, et al.
Published: (2025)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
by: Ren, Feng, et al.
Published: (2026)
by: Ren, Feng, et al.
Published: (2026)
Tetris: Efficient Intra-Datacenter Calls Packing for Large Conferencing Services
by: Gandhi, Rohan, et al.
Published: (2025)
by: Gandhi, Rohan, et al.
Published: (2025)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
by: Du, Kuntai, et al.
Published: (2025)
by: Du, Kuntai, et al.
Published: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
by: Luo, Jingjia, et al.
Published: (2025)
by: Luo, Jingjia, et al.
Published: (2025)
RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Datacenter Energy Optimized Power Profiles
by: Narayanaswamy, Sreedhar, et al.
Published: (2025)
by: Narayanaswamy, Sreedhar, et al.
Published: (2025)
Serving Compound Inference Systems on Datacenter GPUs
by: Devata, Sriram, et al.
Published: (2026)
by: Devata, Sriram, et al.
Published: (2026)
DCGen 1.1 Technical Report: Generating Datacenter Configurations (including IT, Power, Cooling)
by: Gnibga, Wedan Emmanuel, et al.
Published: (2026)
by: Gnibga, Wedan Emmanuel, et al.
Published: (2026)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
by: Liu, Yunzhao, et al.
Published: (2025)
by: Liu, Yunzhao, et al.
Published: (2025)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
by: Yang, Xinjun, et al.
Published: (2025)
by: Yang, Xinjun, et al.
Published: (2025)
Capsule: Efficient Player Isolation for Datacenters
by: Du, Zhouheng, et al.
Published: (2025)
by: Du, Zhouheng, et al.
Published: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
The Ghost in the Datacenter: Link Flapping, Topology Knowledge Failures, and the FITO Category Mistake
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
by: Zhong, Yinmin, et al.
Published: (2024)
by: Zhong, Yinmin, et al.
Published: (2024)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
by: Xue, Chunyu, et al.
Published: (2026)
by: Xue, Chunyu, et al.
Published: (2026)
LAPS: A Length-Aware-Prefill LLM Serving System
by: She, Jianshu, et al.
Published: (2026)
by: She, Jianshu, et al.
Published: (2026)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
by: Niewenhuis, Dante, et al.
Published: (2026)
by: Niewenhuis, Dante, et al.
Published: (2026)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
by: Chen, Xing, et al.
Published: (2025)
by: Chen, Xing, et al.
Published: (2025)
Uncertainty-Aware Decarbonization for Datacenters
by: Li, Amy, et al.
Published: (2024)
by: Li, Amy, et al.
Published: (2024)
Hotspot-Aware Scheduling of Virtual Machines with Overcommitment for Ultimate Utilization in Cloud Datacenters
by: Wu, Jiaxi, et al.
Published: (2026)
by: Wu, Jiaxi, et al.
Published: (2026)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
by: Hsieh, Chia-chi, et al.
Published: (2026)
by: Hsieh, Chia-chi, et al.
Published: (2026)
Distribution and Management of Datacenter Load Decoupling
by: Lin, Liuzixuan, et al.
Published: (2025)
by: Lin, Liuzixuan, et al.
Published: (2025)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Adaptive, Efficient and Fair Resource Allocation in Cloud Datacenters leveraging Weighted A3C Deep Reinforcement Learning
by: Kumari, Suchi, et al.
Published: (2025)
by: Kumari, Suchi, et al.
Published: (2025)
The Cloud Next Door: Investigating the Environmental and Socioeconomic Strain of Datacenters on Local Communities
by: Ngata, Wacuka, et al.
Published: (2025)
by: Ngata, Wacuka, et al.
Published: (2025)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
by: Chen, Shaoyuan, et al.
Published: (2024)
by: Chen, Shaoyuan, et al.
Published: (2024)
QAOA in Quantum Datacenters: Parallelization, Simulation, and Orchestration
by: Liaqat, Amana, et al.
Published: (2025)
by: Liaqat, Amana, et al.
Published: (2025)
Coordinated Cooling and Compute Management for AI Datacenters
by: Abera, Nardos Belay, et al.
Published: (2026)
by: Abera, Nardos Belay, et al.
Published: (2026)
Characterization of Large Language Model Development in the Datacenter
by: Hu, Qinghao, et al.
Published: (2024)
by: Hu, Qinghao, et al.
Published: (2024)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
Similar Items
-
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
by: Qin, Ruoyu, et al.
Published: (2024) -
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
by: Qin, Ruoyu, et al.
Published: (2025) -
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
by: Wang, Jing, et al.
Published: (2025) -
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
by: Wang, Jiahao, et al.
Published: (2025) -
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)