DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Ruan, Chaoyi, Chen, Yinhe, Tian, Dongqi, Shi, Yandong, Wu, Yongji, Li, Jialin, Li, Cheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Reaching Agreement Among Reasoning LLM Agents
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
di: Xu, Jiale, et al.
Pubblicazione: (2025)
di: Xu, Jiale, et al.
Pubblicazione: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
di: Bai, Fan, et al.
Pubblicazione: (2026)
di: Bai, Fan, et al.
Pubblicazione: (2026)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
di: Ren, Feng, et al.
Pubblicazione: (2026)
di: Ren, Feng, et al.
Pubblicazione: (2026)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024)
di: He, Yiyuan, et al.
Pubblicazione: (2024)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
di: Jin, Yibo, et al.
Pubblicazione: (2024)
di: Jin, Yibo, et al.
Pubblicazione: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
Enabling Elastic Model Serving with MultiWorld
di: Lee, Myungjin, et al.
Pubblicazione: (2024)
di: Lee, Myungjin, et al.
Pubblicazione: (2024)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
di: Qiao, Yifan, et al.
Pubblicazione: (2024)
di: Qiao, Yifan, et al.
Pubblicazione: (2024)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
di: Liang, Yunkai, et al.
Pubblicazione: (2025)
di: Liang, Yunkai, et al.
Pubblicazione: (2025)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
di: Hong, Ke, et al.
Pubblicazione: (2025)
di: Hong, Ke, et al.
Pubblicazione: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024) -
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025) -
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025) -
Reaching Agreement Among Reasoning LLM Agents
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025) -
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)