SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Yinghao, Lan, Tingfeng, Huang, Xiuqi, Lu, Hui, Chen, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
by: Chen, Siyuan, et al.
Published: (2025)
by: Chen, Siyuan, et al.
Published: (2025)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025)
by: Gao, Shihong, et al.
Published: (2025)
Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
RW-TTT: Batched Serving for Request-Owned Test-Time Training State
by: Yang, Jian, et al.
Published: (2026)
by: Yang, Jian, et al.
Published: (2026)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
by: Gao, Luyao, et al.
Published: (2025)
by: Gao, Luyao, et al.
Published: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Fast Heterogeneous Serving: Scalable Mixed-Scale LLM Allocation for SLO-Constrained Inference
by: Cheng, Jiaming, et al.
Published: (2026)
by: Cheng, Jiaming, et al.
Published: (2026)
RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models
by: Cong, Xing, et al.
Published: (2026)
by: Cong, Xing, et al.
Published: (2026)
Tuning the Right Foundation Models is What you Need for Partial Label Learning
by: He, Kuang, et al.
Published: (2025)
by: He, Kuang, et al.
Published: (2025)
IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
by: Tang, Yinghao, et al.
Published: (2026)
by: Tang, Yinghao, et al.
Published: (2026)
Right on Time: Revising Time Series Models by Constraining their Explanations
by: Kraus, Maurice, et al.
Published: (2024)
by: Kraus, Maurice, et al.
Published: (2024)
RLBenchNet: The Right Network for the Right Reinforcement Learning Task
by: Smirnov, Ivan, et al.
Published: (2025)
by: Smirnov, Ivan, et al.
Published: (2025)
Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
by: Zeng, Hui, et al.
Published: (2025)
by: Zeng, Hui, et al.
Published: (2025)
EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving
by: Palladino, Vittorio, et al.
Published: (2026)
by: Palladino, Vittorio, et al.
Published: (2026)
Right Reward Right Time for Federated Learning
by: Nguyen, Thanh Linh, et al.
Published: (2025)
by: Nguyen, Thanh Linh, et al.
Published: (2025)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
by: Li, Ruixiao, et al.
Published: (2025)
by: Li, Ruixiao, et al.
Published: (2025)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
by: Sun, Ting, et al.
Published: (2025)
by: Sun, Ting, et al.
Published: (2025)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
by: Lu, Runyu, et al.
Published: (2025)
by: Lu, Runyu, et al.
Published: (2025)
LLMs as Assessors: Right for the Right Reason?
by: Saha, Sourav, et al.
Published: (2026)
by: Saha, Sourav, et al.
Published: (2026)
Right for the Right Reasons: Avoiding Reasoning Shortcuts via Prototypical Neurosymbolic AI
by: Andolfi, Luca, et al.
Published: (2025)
by: Andolfi, Luca, et al.
Published: (2025)
Exploring Fairness in Educational Data Mining in the Context of the Right to be Forgotten
by: Qian, Wei, et al.
Published: (2024)
by: Qian, Wei, et al.
Published: (2024)
Selecting the Right LLM for eGov Explanations
by: Limonad, Lior, et al.
Published: (2025)
by: Limonad, Lior, et al.
Published: (2025)
An Efficient and Generalizable Symbolic Regression Method for Time Series Analysis
by: Xie, Yi, et al.
Published: (2024)
by: Xie, Yi, et al.
Published: (2024)
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
by: Staufer, Dimitri
Published: (2025)
by: Staufer, Dimitri
Published: (2025)
Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
by: Zhuang, Zhan, et al.
Published: (2025)
by: Zhuang, Zhan, et al.
Published: (2025)
Joint Time Series Chain: Detecting Unusual Evolving Trend across Time Series
by: Zhang, Li, et al.
Published: (2026)
by: Zhang, Li, et al.
Published: (2026)
Requests of a Feather Must Flock Together: Batch Size vs. Prefix Homogeneity in LLM Inference
by: Rathi, Saksham, et al.
Published: (2026)
by: Rathi, Saksham, et al.
Published: (2026)
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
by: He, Bingxiang, et al.
Published: (2024)
by: He, Bingxiang, et al.
Published: (2024)
Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
by: Zhang, Qingyang, et al.
Published: (2025)
by: Zhang, Qingyang, et al.
Published: (2025)
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
by: Razavi, Kamran, et al.
Published: (2024)
by: Razavi, Kamran, et al.
Published: (2024)
HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
by: Hu, Zhengding, et al.
Published: (2025)
by: Hu, Zhengding, et al.
Published: (2025)
Finer is Better (with the Right Scaling)
by: Schaefer, Clemens, et al.
Published: (2026)
by: Schaefer, Clemens, et al.
Published: (2026)
TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications
by: Ling, Neiwen, et al.
Published: (2024)
by: Ling, Neiwen, et al.
Published: (2024)
Test-Time Training Done Right
by: Zhang, Tianyuan, et al.
Published: (2025)
by: Zhang, Tianyuan, et al.
Published: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025)
by: Goel, Kanishk, et al.
Published: (2025)
Similar Items
-
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
by: Chen, Siyuan, et al.
Published: (2025) -
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025) -
Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
by: Wang, Yuhang, et al.
Published: (2025) -
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025) -
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
by: Agrawal, Amey, et al.
Published: (2024)