JITServe: SLO-aware LLM Serving with Imprecise Request Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Wei, Wu, Zhiyu, Mu, Yi, Ning, Rui, Liu, Banruo, Sarda, Nikhil, Lee, Myungjin, Lai, Fan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
Enabling Elastic Model Serving with MultiWorld
von: Lee, Myungjin, et al.
Veröffentlicht: (2024)
von: Lee, Myungjin, et al.
Veröffentlicht: (2024)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
von: Nian, Sean, et al.
Veröffentlicht: (2026)
von: Nian, Sean, et al.
Veröffentlicht: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services
von: Liu, Jiachen, et al.
Veröffentlicht: (2024)
von: Liu, Jiachen, et al.
Veröffentlicht: (2024)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
von: Hu, Bodun, et al.
Veröffentlicht: (2024)
von: Hu, Bodun, et al.
Veröffentlicht: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
von: Sun, Desen, et al.
Veröffentlicht: (2025)
von: Sun, Desen, et al.
Veröffentlicht: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
von: Xuan, Mo, et al.
Veröffentlicht: (2025)
von: Xuan, Mo, et al.
Veröffentlicht: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Software-Defined Agentic Serving
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2026)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2026)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
von: Ning, Rui, et al.
Veröffentlicht: (2026)
von: Ning, Rui, et al.
Veröffentlicht: (2026)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
von: Chow, Will
Veröffentlicht: (2025)
von: Chow, Will
Veröffentlicht: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Energy-aware Distributed Microservice Request Placement at the Edge
von: Toczé, Klervie, et al.
Veröffentlicht: (2024)
von: Toczé, Klervie, et al.
Veröffentlicht: (2024)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
Locality-aware Fair Scheduling in LLM Serving
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
von: Zhong, Xinrui, et al.
Veröffentlicht: (2025)
von: Zhong, Xinrui, et al.
Veröffentlicht: (2025)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026) -
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024) -
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024) -
Enabling Elastic Model Serving with MultiWorld
von: Lee, Myungjin, et al.
Veröffentlicht: (2024) -
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
von: Nian, Sean, et al.
Veröffentlicht: (2026)