Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Yukang, Cui, Weihao, Zhao, Han, Xu, Ziyi, Fan, Xiaoze, Chen, Xusheng, Zhou, Yangjie, Sun, Shixuan, He, Bingsheng, Chen, Quan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Function-as-a-Service for Large Language Models with TIDAL
di: Cui, Weihao, et al.
Pubblicazione: (2025)
di: Cui, Weihao, et al.
Pubblicazione: (2025)
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
di: Cui, Weihao, et al.
Pubblicazione: (2025)
di: Cui, Weihao, et al.
Pubblicazione: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
di: Qiu, Shi, et al.
Pubblicazione: (2026)
di: Qiu, Shi, et al.
Pubblicazione: (2026)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
di: Xu, Yuhang, et al.
Pubblicazione: (2025)
di: Xu, Yuhang, et al.
Pubblicazione: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
di: Wu, Yinpeng, et al.
Pubblicazione: (2026)
di: Wu, Yinpeng, et al.
Pubblicazione: (2026)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
di: Zou, Jing, et al.
Pubblicazione: (2026)
di: Zou, Jing, et al.
Pubblicazione: (2026)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
di: Tan, Boyu, et al.
Pubblicazione: (2026)
di: Tan, Boyu, et al.
Pubblicazione: (2026)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
di: Li, Ruihao, et al.
Pubblicazione: (2025)
di: Li, Ruihao, et al.
Pubblicazione: (2025)
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
di: Shen, Weihang, et al.
Pubblicazione: (2025)
di: Shen, Weihang, et al.
Pubblicazione: (2025)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
di: Luo, Shutian, et al.
Pubblicazione: (2026)
di: Luo, Shutian, et al.
Pubblicazione: (2026)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
di: Song, Yixin, et al.
Pubblicazione: (2023)
di: Song, Yixin, et al.
Pubblicazione: (2023)
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
di: Xu, Yechen, et al.
Pubblicazione: (2026)
di: Xu, Yechen, et al.
Pubblicazione: (2026)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
di: Zhang, Keyao, et al.
Pubblicazione: (2025)
di: Zhang, Keyao, et al.
Pubblicazione: (2025)
A Task Equalization Allocation Algorithm Incorporating Blocking Estimation and Resource Similarity Analysis for Vehicle Control Real-Time Systems
di: Duan, Qianlong, et al.
Pubblicazione: (2025)
di: Duan, Qianlong, et al.
Pubblicazione: (2025)
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
di: Han, Mingcong, et al.
Pubblicazione: (2025)
di: Han, Mingcong, et al.
Pubblicazione: (2025)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
di: Xing, Jiarong, et al.
Pubblicazione: (2025)
di: Xing, Jiarong, et al.
Pubblicazione: (2025)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
di: Wang, Yuan, et al.
Pubblicazione: (2025)
di: Wang, Yuan, et al.
Pubblicazione: (2025)
The inverse-closed subalgebra of $C^{*}(G,A)$
di: Chen, Jianjun
Pubblicazione: (2025)
di: Chen, Jianjun
Pubblicazione: (2025)
HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
di: Yang, Bicheng, et al.
Pubblicazione: (2025)
di: Yang, Bicheng, et al.
Pubblicazione: (2025)
Convergence of the Laws of Non-Hermitian Sums of Projections
di: Zhou, Max Sun
Pubblicazione: (2024)
di: Zhou, Max Sun
Pubblicazione: (2024)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
Beyond Money: Incentive Effects of Tokenized Ownership on User Contribution in DAOs
di: Kun Chen, et al.
Pubblicazione: (2025)
di: Kun Chen, et al.
Pubblicazione: (2025)
SSV: Sparse Speculative Verification for Efficient LLM Inference
di: Wang, Zhibin, et al.
Pubblicazione: (2026)
di: Wang, Zhibin, et al.
Pubblicazione: (2026)
Boosting File Systems Elegantly: A Transparent NVM Write-ahead Log for Disk File Systems
di: Wang, Guoyu, et al.
Pubblicazione: (2024)
di: Wang, Guoyu, et al.
Pubblicazione: (2024)
Will an electric vehicle manufacturer benefit from sharing services through price leadership strategy and improving the driving range?
di: Ying Wang, et al.
Pubblicazione: (2024)
di: Ying Wang, et al.
Pubblicazione: (2024)
CHRONOS: Compensating Hardware Related Overheads with Native Multi Timer Support for Real-Time Operating Systems
di: Heider, Kay, et al.
Pubblicazione: (2025)
di: Heider, Kay, et al.
Pubblicazione: (2025)
The Primitive Ideal Space of $C(X) \rtimes \mathbb{N}$
di: Chen, Xiaohui, et al.
Pubblicazione: (2025)
di: Chen, Xiaohui, et al.
Pubblicazione: (2025)
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
di: Zhang, Dingyan, et al.
Pubblicazione: (2026)
di: Zhang, Dingyan, et al.
Pubblicazione: (2026)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
di: Prabhu, Ramya, et al.
Pubblicazione: (2024)
di: Prabhu, Ramya, et al.
Pubblicazione: (2024)
Locking in overseas buyers amid geopolitical conflicts
di: Di Fan, et al.
Pubblicazione: (2024)
di: Di Fan, et al.
Pubblicazione: (2024)
The last mile delivery strategy: relay station service or delivery at home?
di: Ruiqi Zhou, et al.
Pubblicazione: (2025)
di: Ruiqi Zhou, et al.
Pubblicazione: (2025)
The Brown Measure of Non-Hermitian Sums of Projections
di: Zhou, Max Sun
Pubblicazione: (2024)
di: Zhou, Max Sun
Pubblicazione: (2024)
Quaternionic Green's Function and the Brown Measure of Atomic Operators
di: Zhou, Max Sun
Pubblicazione: (2024)
di: Zhou, Max Sun
Pubblicazione: (2024)
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
di: Ahmad, Noaman, et al.
Pubblicazione: (2024)
di: Ahmad, Noaman, et al.
Pubblicazione: (2024)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
di: Gao, Lei, et al.
Pubblicazione: (2025)
di: Gao, Lei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Function-as-a-Service for Large Language Models with TIDAL
di: Cui, Weihao, et al.
Pubblicazione: (2025) -
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
di: Cui, Weihao, et al.
Pubblicazione: (2025) -
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025) -
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
di: Qiu, Shi, et al.
Pubblicazione: (2026) -
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
di: Xu, Yuhang, et al.
Pubblicazione: (2025)