Saved in:
| Main Authors: | Yu, Jiahuan, Hu, Mingtao, Lin, Zichao, Zhang, Minjia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.20309 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
by: Yu, Jiahuan, et al.
Published: (2025)
by: Yu, Jiahuan, et al.
Published: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
by: Huang, Shaoyuan, et al.
Published: (2026)
by: Huang, Shaoyuan, et al.
Published: (2026)
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
by: Qi, Sirui, et al.
Published: (2024)
by: Qi, Sirui, et al.
Published: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
by: Wang, Qipeng
Published: (2026)
by: Wang, Qipeng
Published: (2026)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
by: Li, Yufei, et al.
Published: (2025)
by: Li, Yufei, et al.
Published: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025)
by: Tao, Yiheng, et al.
Published: (2025)
Fairness-Aware Job Scheduling for Multi-Job Federated Learning
by: Shi, Yuxin, et al.
Published: (2024)
by: Shi, Yuxin, et al.
Published: (2024)
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
by: Moore, Hayden, et al.
Published: (2025)
by: Moore, Hayden, et al.
Published: (2025)
TRAIL: Trust-Aware Client Scheduling for Semi-Decentralized Federated Learning
by: Hu, Gangqiang, et al.
Published: (2024)
by: Hu, Gangqiang, et al.
Published: (2024)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
by: Li, Luchang, et al.
Published: (2026)
by: Li, Luchang, et al.
Published: (2026)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
by: Xu, Ao, et al.
Published: (2025)
by: Xu, Ao, et al.
Published: (2025)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
by: Levine, Reese, et al.
Published: (2026)
by: Levine, Reese, et al.
Published: (2026)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
by: Yang, Xinjun, et al.
Published: (2025)
by: Yang, Xinjun, et al.
Published: (2025)
Justitia: Fair and Efficient Scheduling of Task-parallel LLM Agents with Selective Pampering
by: Yang, Mingyan, et al.
Published: (2025)
by: Yang, Mingyan, et al.
Published: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
by: Yu, Donglin
Published: (2026)
by: Yu, Donglin
Published: (2026)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
by: Yousefijamarani, Zahra, et al.
Published: (2025)
by: Yousefijamarani, Zahra, et al.
Published: (2025)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
Application of Machine Learning Optimization in Cloud Computing Resource Scheduling and Management
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024)
by: Yang, Hanmei, et al.
Published: (2024)
ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
by: Choi, Seungbeom, et al.
Published: (2025)
by: Choi, Seungbeom, et al.
Published: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
by: Gond, Raja, et al.
Published: (2026)
by: Gond, Raja, et al.
Published: (2026)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
by: Lyu, Hongtao, et al.
Published: (2025)
by: Lyu, Hongtao, et al.
Published: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
by: Shen, Zixu, et al.
Published: (2025)
by: Shen, Zixu, et al.
Published: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025)
by: Goel, Kanishk, et al.
Published: (2025)
On Evaluating Performance of LLM Inference Serving Systems
by: Agrawal, Amey, et al.
Published: (2025)
by: Agrawal, Amey, et al.
Published: (2025)
Lodestar: An Online-Learning LLM Inference Router
by: Lim, Gangmuk, et al.
Published: (2026)
by: Lim, Gangmuk, et al.
Published: (2026)
Similar Items
-
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025) -
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
by: Yu, Jiahuan, et al.
Published: (2025) -
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
by: Huang, Shaoyuan, et al.
Published: (2026) -
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
by: Qi, Sirui, et al.
Published: (2024) -
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)