AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Haiying, Sen, Tanmoy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
von: Huang, Kaiyu, et al.
Veröffentlicht: (2025)
von: Huang, Kaiyu, et al.
Veröffentlicht: (2025)
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
von: Zhang, Genghan, et al.
Veröffentlicht: (2025)
von: Zhang, Genghan, et al.
Veröffentlicht: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Towards Pareto Optimal Throughput in Small Language Model Serving
von: Recasens, Pol G., et al.
Veröffentlicht: (2024)
von: Recasens, Pol G., et al.
Veröffentlicht: (2024)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
Pie: A Programmable Serving System for Emerging LLM Applications
von: Gim, In, et al.
Veröffentlicht: (2025)
von: Gim, In, et al.
Veröffentlicht: (2025)
Multi-Bin Batching for Increasing LLM Inference Throughput
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
Fast Heterogeneous Serving: Scalable Mixed-Scale LLM Allocation for SLO-Constrained Inference
von: Cheng, Jiaming, et al.
Veröffentlicht: (2026)
von: Cheng, Jiaming, et al.
Veröffentlicht: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
von: Lysenstøen, Christian
Veröffentlicht: (2026)
von: Lysenstøen, Christian
Veröffentlicht: (2026)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
von: Du, Yin, et al.
Veröffentlicht: (2026)
von: Du, Yin, et al.
Veröffentlicht: (2026)
Hierarchical Verification of Speculative Beams for Accelerating LLM Inference
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
von: Wang, Ying, et al.
Veröffentlicht: (2025)
von: Wang, Ying, et al.
Veröffentlicht: (2025)
Taming the Titans: A Survey of Efficient LLM Inference Serving
von: Zhen, Ranran, et al.
Veröffentlicht: (2025)
von: Zhen, Ranran, et al.
Veröffentlicht: (2025)
Ensuring Fair LLM Serving Amid Diverse Applications
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference
von: Yang, Hantao, et al.
Veröffentlicht: (2025)
von: Yang, Hantao, et al.
Veröffentlicht: (2025)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
von: Hajra, Suvadeep, et al.
Veröffentlicht: (2026)
von: Hajra, Suvadeep, et al.
Veröffentlicht: (2026)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024)
von: He, Jiaao, et al.
Veröffentlicht: (2024)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
von: Agarwal, Siddhant, et al.
Veröffentlicht: (2024)
von: Agarwal, Siddhant, et al.
Veröffentlicht: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024) -
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
von: Shen, Haiying, et al.
Veröffentlicht: (2025) -
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024) -
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025) -
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
von: Li, Yufei, et al.
Veröffentlicht: (2025)