Learned Best-Effort LLM Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jha, Siddharth, Hooper, Coleman, Liu, Xiaoxuan, Kim, Sehoon, Keutzer, Kurt |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
AI and Memory Wall
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
Taming the Titans: A Survey of Efficient LLM Inference Serving
von: Zhen, Ranran, et al.
Veröffentlicht: (2025)
von: Zhen, Ranran, et al.
Veröffentlicht: (2025)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)
von: Yang, Shang, et al.
Veröffentlicht: (2025)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
von: Xu, Tianhao, et al.
Veröffentlicht: (2026)
von: Xu, Tianhao, et al.
Veröffentlicht: (2026)
Synera: Synergistic LLM Serving across Device and Cloud at Scale
von: Wang, Genglin, et al.
Veröffentlicht: (2025)
von: Wang, Genglin, et al.
Veröffentlicht: (2025)
Compliance-Scored Best-of-N Guardrail Orchestration for Multimodal Document Generation in Payments Dispute Defense
von: Sundar, Nataraj Agaram, et al.
Veröffentlicht: (2026)
von: Sundar, Nataraj Agaram, et al.
Veröffentlicht: (2026)
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
von: Qiu, Haoran, et al.
Veröffentlicht: (2024)
von: Qiu, Haoran, et al.
Veröffentlicht: (2024)
Autellix: An Efficient Serving Engine for LLM Agents as General Programs
von: Luo, Michael, et al.
Veröffentlicht: (2025)
von: Luo, Michael, et al.
Veröffentlicht: (2025)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
von: Yu, Jiahuan, et al.
Veröffentlicht: (2025)
von: Yu, Jiahuan, et al.
Veröffentlicht: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
von: Hu, Huanqi, et al.
Veröffentlicht: (2025)
von: Hu, Huanqi, et al.
Veröffentlicht: (2025)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2025)
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2025)
Liger Kernel: Efficient Triton Kernels for LLM Training
von: Hsu, Pin-Lun, et al.
Veröffentlicht: (2024)
von: Hsu, Pin-Lun, et al.
Veröffentlicht: (2024)
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
von: Liang, Wanchao, et al.
Veröffentlicht: (2024)
von: Liang, Wanchao, et al.
Veröffentlicht: (2024)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
von: Li, Allison, et al.
Veröffentlicht: (2025)
von: Li, Allison, et al.
Veröffentlicht: (2025)
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
von: Wei, Anjiang, et al.
Veröffentlicht: (2024)
von: Wei, Anjiang, et al.
Veröffentlicht: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023) -
AI and Memory Wall
von: Gholami, Amir, et al.
Veröffentlicht: (2024) -
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
von: Guo, Yipin, et al.
Veröffentlicht: (2026) -
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025) -
Taming the Titans: A Survey of Efficient LLM Inference Serving
von: Zhen, Ranran, et al.
Veröffentlicht: (2025)