Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Yoshimura, Takeshi, van de Beek, Valentijn Dymphnus, Chiba, Tatsuhiro |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speeding up Model Loading with fastsafetensors
di: Yoshimura, Takeshi, et al.
Pubblicazione: (2025)
di: Yoshimura, Takeshi, et al.
Pubblicazione: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
di: Xie, Zhiqiang, et al.
Pubblicazione: (2025)
di: Xie, Zhiqiang, et al.
Pubblicazione: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
di: Bu, Tianci, et al.
Pubblicazione: (2026)
di: Bu, Tianci, et al.
Pubblicazione: (2026)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
di: Li, Cong, et al.
Pubblicazione: (2025)
di: Li, Cong, et al.
Pubblicazione: (2025)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
di: Peng, You, et al.
Pubblicazione: (2026)
di: Peng, You, et al.
Pubblicazione: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
di: Yuan, Ying, et al.
Pubblicazione: (2026)
di: Yuan, Ying, et al.
Pubblicazione: (2026)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
di: Dai, Yinwei, et al.
Pubblicazione: (2025)
di: Dai, Yinwei, et al.
Pubblicazione: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
di: Liu, Di, et al.
Pubblicazione: (2026)
di: Liu, Di, et al.
Pubblicazione: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
di: Shen, Haiying, et al.
Pubblicazione: (2024)
di: Shen, Haiying, et al.
Pubblicazione: (2024)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
Predictable LLM Serving on GPU Clusters
di: Darzi, Erfan, et al.
Pubblicazione: (2025)
di: Darzi, Erfan, et al.
Pubblicazione: (2025)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Fast State Restoration in LLM Serving with HCache
di: Gao, Shiwei, et al.
Pubblicazione: (2024)
di: Gao, Shiwei, et al.
Pubblicazione: (2024)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
di: Xu, Jiale, et al.
Pubblicazione: (2025)
di: Xu, Jiale, et al.
Pubblicazione: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
di: Wu, Bingyang, et al.
Pubblicazione: (2024)
di: Wu, Bingyang, et al.
Pubblicazione: (2024)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
di: Zhang, Chen, et al.
Pubblicazione: (2025)
di: Zhang, Chen, et al.
Pubblicazione: (2025)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Speeding up Model Loading with fastsafetensors
di: Yoshimura, Takeshi, et al.
Pubblicazione: (2025) -
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025) -
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024) -
Strata: Hierarchical Context Caching for Long Context Language Model Serving
di: Xie, Zhiqiang, et al.
Pubblicazione: (2025) -
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)