λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yu, Minchen, Yang, Rui, Jia, Chaobo, Su, Zhaoyuan, Yao, Sheng, Lan, Tingfeng, Yang, Yuchen, Wang, Zirui, Cheng, Yue, Wang, Wei, Wang, Ao, Chen, Ruichuan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
par: Yu, Minchen, et autres
Publié: (2023)
par: Yu, Minchen, et autres
Publié: (2023)
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
par: Wang, Zirui, et autres
Publié: (2025)
par: Wang, Zirui, et autres
Publié: (2025)
LatentBox: Storing AI-Generated Images at Scale via a Latent-First Design
par: Wang, Zirui, et autres
Publié: (2026)
par: Wang, Zirui, et autres
Publié: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
par: Su, Zhaoyuan, et autres
Publié: (2025)
par: Su, Zhaoyuan, et autres
Publié: (2025)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
par: Li, Rui, et autres
Publié: (2025)
par: Li, Rui, et autres
Publié: (2025)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
par: Yu, Minchen, et autres
Publié: (2025)
par: Yu, Minchen, et autres
Publié: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
par: Wu, Hao, et autres
Publié: (2024)
par: Wu, Hao, et autres
Publié: (2024)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
par: Liu, Chongpeng, et autres
Publié: (2025)
par: Liu, Chongpeng, et autres
Publié: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
par: Zhang, Zhexiang, et autres
Publié: (2025)
par: Zhang, Zhexiang, et autres
Publié: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
par: Hu, Junhao, et autres
Publié: (2025)
par: Hu, Junhao, et autres
Publié: (2025)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
par: Ni, Yinan, et autres
Publié: (2025)
par: Ni, Yinan, et autres
Publié: (2025)
TStore: Rethinking AI Model Hub with Tensor-Centric Compression
par: Lan, Tingfeng, et autres
Publié: (2026)
par: Lan, Tingfeng, et autres
Publié: (2026)
HotSwap: Enabling Live Dependency Sharing in Serverless Computing
par: Li, Rui, et autres
Publié: (2024)
par: Li, Rui, et autres
Publié: (2024)
Barrier-Augmented Lagrangian for GPU-based Elastodynamic Contact
par: Guo, Dewen, et autres
Publié: (2024)
par: Guo, Dewen, et autres
Publié: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
par: Li, Haley, et autres
Publié: (2026)
par: Li, Haley, et autres
Publié: (2026)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
par: Chen, Chen, et autres
Publié: (2026)
par: Chen, Chen, et autres
Publié: (2026)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
par: Tzenetopoulos, Achilleas, et autres
Publié: (2024)
par: Tzenetopoulos, Achilleas, et autres
Publié: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
par: Sui, Yifan, et autres
Publié: (2025)
par: Sui, Yifan, et autres
Publié: (2025)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
par: Razavi, Kamran, et autres
Publié: (2024)
par: Razavi, Kamran, et autres
Publié: (2024)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
par: Duan, Jiaang, et autres
Publié: (2024)
par: Duan, Jiaang, et autres
Publié: (2024)
Litmus: Fair Pricing for Serverless Computing
par: Pei, Qi, et autres
Publié: (2024)
par: Pei, Qi, et autres
Publié: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
par: Xu, Chuhao, et autres
Publié: (2025)
par: Xu, Chuhao, et autres
Publié: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
par: Zhao, Han, et autres
Publié: (2024)
par: Zhao, Han, et autres
Publié: (2024)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
par: Chang, Zihan, et autres
Publié: (2024)
par: Chang, Zihan, et autres
Publié: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
par: Gu, Jianfeng, et autres
Publié: (2025)
par: Gu, Jianfeng, et autres
Publié: (2025)
GraphFlash: Enabling Fast and Elastic Graph Processing on Serverless Infrastructure
par: Zhao, Chen, et autres
Publié: (2026)
par: Zhao, Chen, et autres
Publié: (2026)
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
par: Razavi, Kamran, et autres
Publié: (2024)
par: Razavi, Kamran, et autres
Publié: (2024)
Caching Aided Multi-Tenant Serverless Computing
par: Qiao, Chu, et autres
Publié: (2024)
par: Qiao, Chu, et autres
Publié: (2024)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
par: Wu, Z., et autres
Publié: (2025)
par: Wu, Z., et autres
Publié: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
par: Liu, Mengfan, et autres
Publié: (2025)
par: Liu, Mengfan, et autres
Publié: (2025)
KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless Computing
par: Qi, Sheng, et autres
Publié: (2026)
par: Qi, Sheng, et autres
Publié: (2026)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
par: Lou, Chiheng, et autres
Publié: (2025)
par: Lou, Chiheng, et autres
Publié: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
par: Yang, Zheming, et autres
Publié: (2025)
par: Yang, Zheming, et autres
Publié: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
par: Ghosh, Himel
Publié: (2024)
par: Ghosh, Himel
Publié: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
par: Fu, Yao, et autres
Publié: (2024)
par: Fu, Yao, et autres
Publié: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
par: Li, Suyi, et autres
Publié: (2024)
par: Li, Suyi, et autres
Publié: (2024)
Towards Seamless Serverless Computing Across an Edge-Cloud Continuum
par: Simion, Emilian, et autres
Publié: (2024)
par: Simion, Emilian, et autres
Publié: (2024)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
par: Chen, Jiu, et autres
Publié: (2026)
par: Chen, Jiu, et autres
Publié: (2026)
ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
par: Lan, Tingfeng, et autres
Publié: (2025)
par: Lan, Tingfeng, et autres
Publié: (2025)
Documents similaires
-
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
par: Yu, Minchen, et autres
Publié: (2023) -
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
par: Wang, Zirui, et autres
Publié: (2025) -
LatentBox: Storing AI-Generated Images at Scale via a Latent-First Design
par: Wang, Zirui, et autres
Publié: (2026) -
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
par: Su, Zhaoyuan, et autres
Publié: (2025) -
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
par: Li, Rui, et autres
Publié: (2025)