Preble: Efficient Distributed Prompt Scheduling for LLM Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Srivatsa, Vikranth, He, Zijian, Abhyankar, Reyna, Li, Dongming, Zhang, Yiying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InferCept: Efficient Intercept Support for Augmented Large Language Model Inference
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2024)
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2024)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
Locality-aware Fair Scheduling in LLM Serving
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
von: Yao, Yuhang, et al.
Veröffentlicht: (2024)
von: Yao, Yuhang, et al.
Veröffentlicht: (2024)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
von: Bai, Xu, et al.
Veröffentlicht: (2026)
von: Bai, Xu, et al.
Veröffentlicht: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
von: Lee, Gunjun, et al.
Veröffentlicht: (2025)
von: Lee, Gunjun, et al.
Veröffentlicht: (2025)
Prompt-Aware Scheduling for Efficient Text-to-Image Inferencing System
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
Llumnix: Dynamic Scheduling for Large Language Model Serving
von: Sun, Biao, et al.
Veröffentlicht: (2024)
von: Sun, Biao, et al.
Veröffentlicht: (2024)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels
von: Song, Mingcong, et al.
Veröffentlicht: (2024)
von: Song, Mingcong, et al.
Veröffentlicht: (2024)
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
InferCept: Efficient Intercept Support for Augmented Large Language Model Inference
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2024) -
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026) -
Locality-aware Fair Scheduling in LLM Serving
von: Cao, Shiyi, et al.
Veröffentlicht: (2025) -
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025) -
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)