Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Mingcong, Tang, Xinru, Hou, Fengfan, Li, Jing, Wei, Wei, Ma, Yipeng, Xiao, Runqiu, Si, Hongjie, Jiang, Dingcheng, Yin, Shouyi, Hu, Yang, Long, Guoping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
di: Hsieh, Chia-chi, et al.
Pubblicazione: (2026)
di: Hsieh, Chia-chi, et al.
Pubblicazione: (2026)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
di: Tang, Xinru, et al.
Pubblicazione: (2025)
di: Tang, Xinru, et al.
Pubblicazione: (2025)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
di: Kanani, Alish, et al.
Pubblicazione: (2026)
di: Kanani, Alish, et al.
Pubblicazione: (2026)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
di: She, Jianshu, et al.
Pubblicazione: (2026)
di: She, Jianshu, et al.
Pubblicazione: (2026)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
di: Du, Kuntai, et al.
Pubblicazione: (2025)
di: Du, Kuntai, et al.
Pubblicazione: (2025)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
di: Bu, Tianci, et al.
Pubblicazione: (2026)
di: Bu, Tianci, et al.
Pubblicazione: (2026)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
di: Da, Wei, et al.
Pubblicazione: (2025)
di: Da, Wei, et al.
Pubblicazione: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
di: Pagonas, Nikos, et al.
Pubblicazione: (2025)
di: Pagonas, Nikos, et al.
Pubblicazione: (2025)
Past-Future Scheduler for LLM Serving under SLA Guarantees
di: Gong, Ruihao, et al.
Pubblicazione: (2025)
di: Gong, Ruihao, et al.
Pubblicazione: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
PHWSOA: A Pareto-based Hybrid Whale-Seagull Scheduling for Multi-Objective Tasks in Cloud Computing
di: Zhao, Zhi, et al.
Pubblicazione: (2025)
di: Zhao, Zhi, et al.
Pubblicazione: (2025)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
di: He, Xuan, et al.
Pubblicazione: (2025)
di: He, Xuan, et al.
Pubblicazione: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
di: Wolfrath, Joel, et al.
Pubblicazione: (2025)
di: Wolfrath, Joel, et al.
Pubblicazione: (2025)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
di: Zhao, Zhixin, et al.
Pubblicazione: (2024)
di: Zhao, Zhixin, et al.
Pubblicazione: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
di: Peng, You, et al.
Pubblicazione: (2026)
di: Peng, You, et al.
Pubblicazione: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
di: Zhou, Bowen, et al.
Pubblicazione: (2026)
di: Zhou, Bowen, et al.
Pubblicazione: (2026)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
Large-Scale LLM Inference with Heterogeneous Workloads: Prefill-Decode Contention and Asymptotically Optimal Control
di: Lin, Ruihan, et al.
Pubblicazione: (2026)
di: Lin, Ruihan, et al.
Pubblicazione: (2026)
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
di: Qin, Ruoyu, et al.
Pubblicazione: (2026)
di: Qin, Ruoyu, et al.
Pubblicazione: (2026)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
Locality-aware Fair Scheduling in LLM Serving
di: Cao, Shiyi, et al.
Pubblicazione: (2025)
di: Cao, Shiyi, et al.
Pubblicazione: (2025)
In Serverless, OS Scheduler Choice Costs Money: A Hybrid Scheduling Approach for Cheaper FaaS
di: Zhao, Yuxuan, et al.
Pubblicazione: (2024)
di: Zhao, Yuxuan, et al.
Pubblicazione: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
di: Yang, Qizheng, et al.
Pubblicazione: (2025)
di: Yang, Qizheng, et al.
Pubblicazione: (2025)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
di: Zhang, Hengrui, et al.
Pubblicazione: (2025)
di: Zhang, Hengrui, et al.
Pubblicazione: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024) -
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025) -
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025) -
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
di: Hsieh, Chia-chi, et al.
Pubblicazione: (2026) -
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)