TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
Fuente:
arXiv
Salvato in:
| Autori principali: | Lai, Ruiqi, Liu, Hongrui, Lu, Chengzhi, Liu, Zonghao, Cao, Siyu, Shao, Siyang, Zhang, Yixin, Mai, Luo, Ustiugov, Dmitrii |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
di: Kondrashov, Leonid, et al.
Pubblicazione: (2025)
di: Kondrashov, Leonid, et al.
Pubblicazione: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
di: Gao, Wei, et al.
Pubblicazione: (2026)
di: Gao, Wei, et al.
Pubblicazione: (2026)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
di: Huang, Tao, et al.
Pubblicazione: (2024)
di: Huang, Tao, et al.
Pubblicazione: (2024)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
di: Wu, Yu, et al.
Pubblicazione: (2025)
di: Wu, Yu, et al.
Pubblicazione: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
di: Fu, Yao, et al.
Pubblicazione: (2024)
di: Fu, Yao, et al.
Pubblicazione: (2024)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
di: Ren, Feng, et al.
Pubblicazione: (2026)
di: Ren, Feng, et al.
Pubblicazione: (2026)
OTAS: An Elastic Transformer Serving System via Token Adaptation
di: Chen, Jinyu, et al.
Pubblicazione: (2024)
di: Chen, Jinyu, et al.
Pubblicazione: (2024)
Hierarchical Autoscaling for Large Language Model Serving with Chiron
di: Patke, Archit, et al.
Pubblicazione: (2025)
di: Patke, Archit, et al.
Pubblicazione: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
di: Wang, Shao, et al.
Pubblicazione: (2026)
di: Wang, Shao, et al.
Pubblicazione: (2026)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
Past-Future Scheduler for LLM Serving under SLA Guarantees
di: Gong, Ruihao, et al.
Pubblicazione: (2025)
di: Gong, Ruihao, et al.
Pubblicazione: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
di: Lin, Yanying, et al.
Pubblicazione: (2025)
di: Lin, Yanying, et al.
Pubblicazione: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
di: Jin, Yibo, et al.
Pubblicazione: (2024)
di: Jin, Yibo, et al.
Pubblicazione: (2024)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
di: Wu, Siyu, et al.
Pubblicazione: (2025)
di: Wu, Siyu, et al.
Pubblicazione: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
Self-adaptive, Requirements-driven Autoscaling of Microservices
di: Nunes, João Paulo Karol Santos, et al.
Pubblicazione: (2024)
di: Nunes, João Paulo Karol Santos, et al.
Pubblicazione: (2024)
Proactive and Reactive Autoscaling Techniques for Edge Computing
di: Gupta, Suhrid, et al.
Pubblicazione: (2025)
di: Gupta, Suhrid, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
di: Chen, Wenyan, et al.
Pubblicazione: (2026) -
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
di: Li, Rongzhi, et al.
Pubblicazione: (2025) -
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
di: Kondrashov, Leonid, et al.
Pubblicazione: (2025) -
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026) -
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)