Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
Fuente:
arXiv
Salvato in:
| Autore principale: | Zhou, Zijie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
di: Kamahori, Keisuke, et al.
Pubblicazione: (2026)
di: Kamahori, Keisuke, et al.
Pubblicazione: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
di: Cheng, Rongxin, et al.
Pubblicazione: (2024)
di: Cheng, Rongxin, et al.
Pubblicazione: (2024)
LLM Inference Serving: Survey of Recent Advances and Opportunities
di: Li, Baolin, et al.
Pubblicazione: (2024)
di: Li, Baolin, et al.
Pubblicazione: (2024)
LAPS: A Length-Aware-Prefill LLM Serving System
di: She, Jianshu, et al.
Pubblicazione: (2026)
di: She, Jianshu, et al.
Pubblicazione: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
di: Wagenländer, Marcel, et al.
Pubblicazione: (2026)
di: Wagenländer, Marcel, et al.
Pubblicazione: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
di: Hankendi, Can, et al.
Pubblicazione: (2026)
di: Hankendi, Can, et al.
Pubblicazione: (2026)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
di: Huang, Tao, et al.
Pubblicazione: (2024)
di: Huang, Tao, et al.
Pubblicazione: (2024)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
di: Da, Wei, et al.
Pubblicazione: (2025)
di: Da, Wei, et al.
Pubblicazione: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
di: Yousefijamarani, Zahra, et al.
Pubblicazione: (2025)
di: Yousefijamarani, Zahra, et al.
Pubblicazione: (2025)
TS-EoH: An Edge Server Task Scheduling Algorithm Based on Evolution of Heuristic
di: Yatong, Wang, et al.
Pubblicazione: (2024)
di: Yatong, Wang, et al.
Pubblicazione: (2024)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
di: Wang, Haodong, et al.
Pubblicazione: (2025)
di: Wang, Haodong, et al.
Pubblicazione: (2025)
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
di: Fang, Fei, et al.
Pubblicazione: (2025)
di: Fang, Fei, et al.
Pubblicazione: (2025)
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
di: Xu, Tianhao, et al.
Pubblicazione: (2026)
di: Xu, Tianhao, et al.
Pubblicazione: (2026)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
di: Song, Xiaoniu, et al.
Pubblicazione: (2024)
di: Song, Xiaoniu, et al.
Pubblicazione: (2024)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
di: Cho, Jaehong, et al.
Pubblicazione: (2024)
di: Cho, Jaehong, et al.
Pubblicazione: (2024)
A Parallel CPU-GPU Framework for Batching Heuristic Operations in Depth-First Heuristic Search
di: Futuhi, Ehsan, et al.
Pubblicazione: (2025)
di: Futuhi, Ehsan, et al.
Pubblicazione: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
di: Mao, Ziming, et al.
Pubblicazione: (2024)
di: Mao, Ziming, et al.
Pubblicazione: (2024)
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
di: Cho, Jaehong, et al.
Pubblicazione: (2025)
di: Cho, Jaehong, et al.
Pubblicazione: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
di: Bai, Fan, et al.
Pubblicazione: (2026)
di: Bai, Fan, et al.
Pubblicazione: (2026)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
Regulating Branch Parallelism in LLM Serving
di: Gandhi, Swapnil, et al.
Pubblicazione: (2026)
di: Gandhi, Swapnil, et al.
Pubblicazione: (2026)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
di: Hsieh, Chia-chi, et al.
Pubblicazione: (2026)
di: Hsieh, Chia-chi, et al.
Pubblicazione: (2026)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
di: Yuan, Yichao, et al.
Pubblicazione: (2025)
di: Yuan, Yichao, et al.
Pubblicazione: (2025)
A Meta-Heuristic Load Balancer for Cloud Computing Systems
di: Sliwko, Leszek, et al.
Pubblicazione: (2025)
di: Sliwko, Leszek, et al.
Pubblicazione: (2025)
Slicing Is All You Need: Towards A Universal One-Sided Algorithm for Distributed Matrix Multiplication
di: Brock, Benjamin, et al.
Pubblicazione: (2025)
di: Brock, Benjamin, et al.
Pubblicazione: (2025)
Hierarchical Autoscaling for Large Language Model Serving with Chiron
di: Patke, Archit, et al.
Pubblicazione: (2025)
di: Patke, Archit, et al.
Pubblicazione: (2025)
Synera: Synergistic LLM Serving across Device and Cloud at Scale
di: Wang, Genglin, et al.
Pubblicazione: (2025)
di: Wang, Genglin, et al.
Pubblicazione: (2025)
LegoDiffusion: Micro-Serving Text-to-Image Diffusion Workflows
di: Yang, Lingyun, et al.
Pubblicazione: (2026)
di: Yang, Lingyun, et al.
Pubblicazione: (2026)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
di: Goel, Kanishk, et al.
Pubblicazione: (2025)
di: Goel, Kanishk, et al.
Pubblicazione: (2025)
On Evaluating Performance of LLM Inference Serving Systems
di: Agrawal, Amey, et al.
Pubblicazione: (2025)
di: Agrawal, Amey, et al.
Pubblicazione: (2025)
Tutoring LLM into a Better CUDA Optimizer
di: Brabec, Matyáš, et al.
Pubblicazione: (2025)
di: Brabec, Matyáš, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025) -
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
di: Kamahori, Keisuke, et al.
Pubblicazione: (2026) -
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026) -
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025) -
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
di: Cheng, Rongxin, et al.
Pubblicazione: (2024)