Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
Fuente:
arXiv
Salvato in:
| Autori principali: | Liang, Yunkai, Chen, Zhangyu, Zuo, Pengfei, Zhou, Zhi, Chen, Xu, Yu, Zhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
di: Yuan, Ying, et al.
Pubblicazione: (2026)
di: Yuan, Ying, et al.
Pubblicazione: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
di: Lin, Zejia, et al.
Pubblicazione: (2025)
di: Lin, Zejia, et al.
Pubblicazione: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
di: Masood, Amna, et al.
Pubblicazione: (2026)
di: Masood, Amna, et al.
Pubblicazione: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
di: Shen, Haiying, et al.
Pubblicazione: (2024)
di: Shen, Haiying, et al.
Pubblicazione: (2024)
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
di: Hong, Ke, et al.
Pubblicazione: (2025)
di: Hong, Ke, et al.
Pubblicazione: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
KVDirect: Distributed Disaggregated LLM Inference
di: Chen, Shiyang, et al.
Pubblicazione: (2024)
di: Chen, Shiyang, et al.
Pubblicazione: (2024)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
di: Zhao, Lingxiao, et al.
Pubblicazione: (2025)
di: Zhao, Lingxiao, et al.
Pubblicazione: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
di: Liu, Di, et al.
Pubblicazione: (2026)
di: Liu, Di, et al.
Pubblicazione: (2026)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
di: Chen, Shaoyuan, et al.
Pubblicazione: (2024)
di: Chen, Shaoyuan, et al.
Pubblicazione: (2024)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
di: Jin, Yibo, et al.
Pubblicazione: (2024)
di: Jin, Yibo, et al.
Pubblicazione: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
di: Wang, Shao, et al.
Pubblicazione: (2026)
di: Wang, Shao, et al.
Pubblicazione: (2026)
Efficient Long-context Language Model Training by Core Attention Disaggregation
di: Zhuang, Yonghao, et al.
Pubblicazione: (2025)
di: Zhuang, Yonghao, et al.
Pubblicazione: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
di: Qiao, Yifan, et al.
Pubblicazione: (2024)
di: Qiao, Yifan, et al.
Pubblicazione: (2024)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
di: Guo, Yipin, et al.
Pubblicazione: (2026)
di: Guo, Yipin, et al.
Pubblicazione: (2026)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
di: Zhong, Yinmin, et al.
Pubblicazione: (2025)
di: Zhong, Yinmin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025) -
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
di: Yuan, Ying, et al.
Pubblicazione: (2026) -
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025) -
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025) -
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
di: Lin, Zejia, et al.
Pubblicazione: (2025)