Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Ghosh, Himel |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
par: Fu, Yao, et autres
Publié: (2024)
par: Fu, Yao, et autres
Publié: (2024)
Collaborative Speculative Inference for Efficient LLM Inference Serving
par: Gao, Luyao, et autres
Publié: (2025)
par: Gao, Luyao, et autres
Publié: (2025)
Fast Distributed Inference Serving for Large Language Models
par: Wu, Bingyang, et autres
Publié: (2023)
par: Wu, Bingyang, et autres
Publié: (2023)
DeepServe: Serverless Large Language Model Serving at Scale
par: Hu, Junhao, et autres
Publié: (2025)
par: Hu, Junhao, et autres
Publié: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
par: Yu, Hanfei, et autres
Publié: (2026)
par: Yu, Hanfei, et autres
Publié: (2026)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
par: Kim, Kihyun, et autres
Publié: (2025)
par: Kim, Kihyun, et autres
Publié: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
par: Wu, Bingyang, et autres
Publié: (2024)
par: Wu, Bingyang, et autres
Publié: (2024)
Apodotiko: Enabling Efficient Serverless Federated Learning in Heterogeneous Environments
par: Chadha, Mohak, et autres
Publié: (2024)
par: Chadha, Mohak, et autres
Publié: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
par: Lou, Chiheng, et autres
Publié: (2025)
par: Lou, Chiheng, et autres
Publié: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
par: Kossmann, Ferdi, et autres
Publié: (2024)
par: Kossmann, Ferdi, et autres
Publié: (2024)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
par: Lou, Chiheng, et autres
Publié: (2025)
par: Lou, Chiheng, et autres
Publié: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
par: Barrak, Amine, et autres
Publié: (2025)
par: Barrak, Amine, et autres
Publié: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
par: Griggs, Tyler, et autres
Publié: (2024)
par: Griggs, Tyler, et autres
Publié: (2024)
CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
par: Jin, Hongpeng, et autres
Publié: (2024)
par: Jin, Hongpeng, et autres
Publié: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
par: Sui, Yifan, et autres
Publié: (2025)
par: Sui, Yifan, et autres
Publié: (2025)
Towards Sustainable Large Language Model Serving
par: Nguyen, Sophia, et autres
Publié: (2024)
par: Nguyen, Sophia, et autres
Publié: (2024)
Stateful Large Language Model Serving with Pensieve
par: Yu, Lingfan, et autres
Publié: (2023)
par: Yu, Lingfan, et autres
Publié: (2023)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
par: Oakley, Joe, et autres
Publié: (2024)
par: Oakley, Joe, et autres
Publié: (2024)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
par: Wang, Minghe, et autres
Publié: (2026)
par: Wang, Minghe, et autres
Publié: (2026)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
par: Liu, Mengfan, et autres
Publié: (2025)
par: Liu, Mengfan, et autres
Publié: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
par: Agrawal, Amey, et autres
Publié: (2024)
par: Agrawal, Amey, et autres
Publié: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
par: Yu, Minchen, et autres
Publié: (2025)
par: Yu, Minchen, et autres
Publié: (2025)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
par: Yu, Fengze, et autres
Publié: (2025)
par: Yu, Fengze, et autres
Publié: (2025)
Preble: Efficient Distributed Prompt Scheduling for LLM Serving
par: Srivatsa, Vikranth, et autres
Publié: (2024)
par: Srivatsa, Vikranth, et autres
Publié: (2024)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
par: Gao, Lei, et autres
Publié: (2024)
par: Gao, Lei, et autres
Publié: (2024)
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2024)
par: Xu, Jiale, et autres
Publié: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
par: Su, Zhaoyuan, et autres
Publié: (2025)
par: Su, Zhaoyuan, et autres
Publié: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
par: Xu, Chuhao, et autres
Publié: (2025)
par: Xu, Chuhao, et autres
Publié: (2025)
Shabari: Delayed Decision-Making for Faster and Efficient Serverless Functions
par: Sinha, Prasoon, et autres
Publié: (2024)
par: Sinha, Prasoon, et autres
Publié: (2024)
PolyServe: Efficient Multi-SLO Serving at Scale
par: Zhu, Kan, et autres
Publié: (2025)
par: Zhu, Kan, et autres
Publié: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
par: Ye, Zihao, et autres
Publié: (2025)
par: Ye, Zihao, et autres
Publié: (2025)
A Universal Load Balancing Principle and Its Application to Large Language Model Serving
par: Chen, Zixi, et autres
Publié: (2026)
par: Chen, Zixi, et autres
Publié: (2026)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
par: Jin, Yibo, et autres
Publié: (2024)
par: Jin, Yibo, et autres
Publié: (2024)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
par: Lu, Runyu, et autres
Publié: (2025)
par: Lu, Runyu, et autres
Publié: (2025)
EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
par: Yu, Zhongzhi, et autres
Publié: (2024)
par: Yu, Zhongzhi, et autres
Publié: (2024)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
par: Bai, Xu, et autres
Publié: (2026)
par: Bai, Xu, et autres
Publié: (2026)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
par: Lee, Wonbeom, et autres
Publié: (2024)
par: Lee, Wonbeom, et autres
Publié: (2024)
FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving
par: Shen, Ao, et autres
Publié: (2024)
par: Shen, Ao, et autres
Publié: (2024)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
par: Miao, Xupeng, et autres
Publié: (2023)
par: Miao, Xupeng, et autres
Publié: (2023)
Documents similaires
-
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
par: Fu, Yao, et autres
Publié: (2024) -
Collaborative Speculative Inference for Efficient LLM Inference Serving
par: Gao, Luyao, et autres
Publié: (2025) -
Fast Distributed Inference Serving for Large Language Models
par: Wu, Bingyang, et autres
Publié: (2023) -
DeepServe: Serverless Large Language Model Serving at Scale
par: Hu, Junhao, et autres
Publié: (2025) -
MoEless: Efficient MoE LLM Serving via Serverless Computing
par: Yu, Hanfei, et autres
Publié: (2026)