DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Yinmin, Liu, Shengyu, Chen, Junda, Hu, Jianbo, Zhu, Yibo, Liu, Xuanzhe, Jin, Xin, Zhang, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
von: Basit, Omar, et al.
Veröffentlicht: (2026)
von: Basit, Omar, et al.
Veröffentlicht: (2026)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
von: Liu, Yunzhao, et al.
Veröffentlicht: (2025)
von: Liu, Yunzhao, et al.
Veröffentlicht: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
von: Ren, Feng, et al.
Veröffentlicht: (2026)
von: Ren, Feng, et al.
Veröffentlicht: (2026)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
von: Kanani, Alish, et al.
Veröffentlicht: (2026)
von: Kanani, Alish, et al.
Veröffentlicht: (2026)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025) -
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024) -
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023) -
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024) -
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)