SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yu, Jiahuan, Hu, Mingtao, Lin, Zichao, Zhang, Minjia |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
par: Lian, Xinyu, et autres
Publié: (2025)
par: Lian, Xinyu, et autres
Publié: (2025)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
par: Yu, Jiahuan, et autres
Publié: (2025)
par: Yu, Jiahuan, et autres
Publié: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
par: Huang, Shaoyuan, et autres
Publié: (2026)
par: Huang, Shaoyuan, et autres
Publié: (2026)
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
par: Qi, Sirui, et autres
Publié: (2024)
par: Qi, Sirui, et autres
Publié: (2024)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
par: Ye, Zihao, et autres
Publié: (2025)
par: Ye, Zihao, et autres
Publié: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
par: Li, Zikun, et autres
Publié: (2025)
par: Li, Zikun, et autres
Publié: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
par: Stojkovic, Jovan, et autres
Publié: (2025)
par: Stojkovic, Jovan, et autres
Publié: (2025)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
par: Kakolyris, Andreas Kosmas, et autres
Publié: (2024)
par: Kakolyris, Andreas Kosmas, et autres
Publié: (2024)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
par: Ao, Ruicheng, et autres
Publié: (2025)
par: Ao, Ruicheng, et autres
Publié: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
par: Jiang, Xuanlin, et autres
Publié: (2024)
par: Jiang, Xuanlin, et autres
Publié: (2024)
Fairness-Aware Job Scheduling for Multi-Job Federated Learning
par: Shi, Yuxin, et autres
Publié: (2024)
par: Shi, Yuxin, et autres
Publié: (2024)
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
par: Moore, Hayden, et autres
Publié: (2025)
par: Moore, Hayden, et autres
Publié: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
par: Tao, Yiheng, et autres
Publié: (2025)
par: Tao, Yiheng, et autres
Publié: (2025)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
par: Li, Yufei, et autres
Publié: (2025)
par: Li, Yufei, et autres
Publié: (2025)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
par: Xu, Ao, et autres
Publié: (2025)
par: Xu, Ao, et autres
Publié: (2025)
TRAIL: Trust-Aware Client Scheduling for Semi-Decentralized Federated Learning
par: Hu, Gangqiang, et autres
Publié: (2024)
par: Hu, Gangqiang, et autres
Publié: (2024)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
par: Li, Luchang, et autres
Publié: (2026)
par: Li, Luchang, et autres
Publié: (2026)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
par: Yang, Xinjun, et autres
Publié: (2025)
par: Yang, Xinjun, et autres
Publié: (2025)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
par: Levine, Reese, et autres
Publié: (2026)
par: Levine, Reese, et autres
Publié: (2026)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
par: Yousefijamarani, Zahra, et autres
Publié: (2025)
par: Yousefijamarani, Zahra, et autres
Publié: (2025)
Justitia: Fair and Efficient Scheduling of Task-parallel LLM Agents with Selective Pampering
par: Yang, Mingyan, et autres
Publié: (2025)
par: Yang, Mingyan, et autres
Publié: (2025)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
par: Yu, Donglin
Publié: (2026)
par: Yu, Donglin
Publié: (2026)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
par: Lyu, Hongtao, et autres
Publié: (2025)
par: Lyu, Hongtao, et autres
Publié: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
par: Lee, Kyle, et autres
Publié: (2026)
par: Lee, Kyle, et autres
Publié: (2026)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
par: Cheng, Rongxin, et autres
Publié: (2024)
par: Cheng, Rongxin, et autres
Publié: (2024)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
par: Zhao, Juntao, et autres
Publié: (2024)
par: Zhao, Juntao, et autres
Publié: (2024)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
par: Kim, Joon Ha, et autres
Publié: (2026)
par: Kim, Joon Ha, et autres
Publié: (2026)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
par: Pan, Xinglin, et autres
Publié: (2025)
par: Pan, Xinglin, et autres
Publié: (2025)
Application of Machine Learning Optimization in Cloud Computing Resource Scheduling and Management
par: Zhang, Yifan, et autres
Publié: (2024)
par: Zhang, Yifan, et autres
Publié: (2024)
ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
par: Choi, Seungbeom, et autres
Publié: (2025)
par: Choi, Seungbeom, et autres
Publié: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
par: Shen, Zixu, et autres
Publié: (2025)
par: Shen, Zixu, et autres
Publié: (2025)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
par: Yang, Hanmei, et autres
Publié: (2024)
par: Yang, Hanmei, et autres
Publié: (2024)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
par: Gond, Raja, et autres
Publié: (2026)
par: Gond, Raja, et autres
Publié: (2026)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
par: Li, Yan, et autres
Publié: (2025)
par: Li, Yan, et autres
Publié: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
par: Ma, Chenxiang, et autres
Publié: (2025)
par: Ma, Chenxiang, et autres
Publié: (2025)
Documents similaires
-
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
par: Lian, Xinyu, et autres
Publié: (2025) -
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
par: Yu, Jiahuan, et autres
Publié: (2025) -
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
par: Huang, Shaoyuan, et autres
Publié: (2026) -
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
par: Qi, Sirui, et autres
Publié: (2024) -
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)