ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Yao, Xue, Leyang, Huang, Yeqi, Brabete, Andrei-Octavian, Ustiugov, Dmitrii, Patel, Yuvraj, Mai, Luo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023)
by: Yu, Minchen, et al.
Published: (2023)
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026)
by: Park, JooYoung, et al.
Published: (2026)
Melding the Serverless Control Plane with the Conventional Cluster Manager for Speed and Resource Efficiency
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
by: Sui, Yifan, et al.
Published: (2025)
by: Sui, Yifan, et al.
Published: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
by: Ghosh, Himel
Published: (2024)
by: Ghosh, Himel
Published: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
Cold Start Latency in Serverless Computing: A Systematic Review, Taxonomy, and Future Directions
by: Golec, Muhammed, et al.
Published: (2023)
by: Golec, Muhammed, et al.
Published: (2023)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
by: Duan, Jiaang, et al.
Published: (2024)
by: Duan, Jiaang, et al.
Published: (2024)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
by: Chen, Wenyan, et al.
Published: (2026)
by: Chen, Wenyan, et al.
Published: (2026)
Are Unikernels Ready for Serverless on the Edge?
by: Moebius, Felix, et al.
Published: (2024)
by: Moebius, Felix, et al.
Published: (2024)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
by: Wu, Z., et al.
Published: (2025)
by: Wu, Z., et al.
Published: (2025)
Shattering the Ephemeral Storage Cost Barrier for Data-Intensive Serverless Workflows
by: Ustiugov, Dmitrii, et al.
Published: (2023)
by: Ustiugov, Dmitrii, et al.
Published: (2023)
Multi-Event Triggers for Serverless Computing
by: Carl, Natalie, et al.
Published: (2025)
by: Carl, Natalie, et al.
Published: (2025)
Raptor: Distributed Scheduling for Serverless Functions
by: Exton, Kevin, et al.
Published: (2024)
by: Exton, Kevin, et al.
Published: (2024)
Energy Efficient Scheduling for Serverless Systems
by: Tsenos, Michail, et al.
Published: (2024)
by: Tsenos, Michail, et al.
Published: (2024)
Affinity-aware Serverless Function Scheduling
by: De Palma, Giuseppe, et al.
Published: (2024)
by: De Palma, Giuseppe, et al.
Published: (2024)
Litmus: Fair Pricing for Serverless Computing
by: Pei, Qi, et al.
Published: (2024)
by: Pei, Qi, et al.
Published: (2024)
Serverless Computing: Architecture, Concepts, and Applications
by: Ghorbian, Mohsen, et al.
Published: (2025)
by: Ghorbian, Mohsen, et al.
Published: (2025)
Komet: A Serverless Platform for Low-Earth Orbit Edge Services
by: Pfandzelter, Tobias, et al.
Published: (2024)
by: Pfandzelter, Tobias, et al.
Published: (2024)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
by: Liu, Chongpeng, et al.
Published: (2025)
by: Liu, Chongpeng, et al.
Published: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
by: Lai, Ruiqi, et al.
Published: (2025)
by: Lai, Ruiqi, et al.
Published: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
by: Liu, Wentao, et al.
Published: (2025)
by: Liu, Wentao, et al.
Published: (2025)
FaasMeter: Energy-First Serverless Computing
by: Rehman, Abdul, et al.
Published: (2024)
by: Rehman, Abdul, et al.
Published: (2024)
Serverless Abstractions for Short-Running, Lightweight Streams
by: Carl, Natalie, et al.
Published: (2026)
by: Carl, Natalie, et al.
Published: (2026)
Orchestrating the Execution of Serverless Functions in Hybrid Clouds
by: Peri, Aristotelis, et al.
Published: (2024)
by: Peri, Aristotelis, et al.
Published: (2024)
Konflux: Optimized Function Fusion for Serverless Applications
by: Kowallik, Niklas, et al.
Published: (2026)
by: Kowallik, Niklas, et al.
Published: (2026)
Caching Aided Multi-Tenant Serverless Computing
by: Qiao, Chu, et al.
Published: (2024)
by: Qiao, Chu, et al.
Published: (2024)
Zenix: Efficient Execution of Bulky Serverless Applications
by: Guo, Zhiyuan, et al.
Published: (2022)
by: Guo, Zhiyuan, et al.
Published: (2022)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
On the Complexity of Reachability Properties in Serverless Function Scheduling
by: De Palma, Giuseppe, et al.
Published: (2024)
by: De Palma, Giuseppe, et al.
Published: (2024)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
by: Ni, Yinan, et al.
Published: (2025)
by: Ni, Yinan, et al.
Published: (2025)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
by: Chang, Zihan, et al.
Published: (2024)
by: Chang, Zihan, et al.
Published: (2024)
Dependency-aware Resource Allocation for Serverless Functions at the Edge
by: Baresi, Luciano, et al.
Published: (2023)
by: Baresi, Luciano, et al.
Published: (2023)
Similar Items
-
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023) -
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
by: Kondrashov, Leonid, et al.
Published: (2025) -
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026) -
Melding the Serverless Control Plane with the Conventional Cluster Manager for Speed and Resource Efficiency
by: Kondrashov, Leonid, et al.
Published: (2025) -
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
by: Sui, Yifan, et al.
Published: (2025)