Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Z., Deng, Y., Hu, J., Cui, L., Zhang, Z., Zeng, L., Min, G. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
di: Kong, Z. Jonny, et al.
Pubblicazione: (2025)
di: Kong, Z. Jonny, et al.
Pubblicazione: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
di: Hui, Xinning, et al.
Pubblicazione: (2024)
di: Hui, Xinning, et al.
Pubblicazione: (2024)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
di: Yu, Minchen, et al.
Pubblicazione: (2023)
di: Yu, Minchen, et al.
Pubblicazione: (2023)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
di: Liu, Chongpeng, et al.
Pubblicazione: (2025)
di: Liu, Chongpeng, et al.
Pubblicazione: (2025)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
di: Duan, Jiaang, et al.
Pubblicazione: (2024)
di: Duan, Jiaang, et al.
Pubblicazione: (2024)
Towards Fast Setup and High Throughput of GPU Serverless Computing
di: Zhao, Han, et al.
Pubblicazione: (2024)
di: Zhao, Han, et al.
Pubblicazione: (2024)
Zenix: Efficient Execution of Bulky Serverless Applications
di: Guo, Zhiyuan, et al.
Pubblicazione: (2022)
di: Guo, Zhiyuan, et al.
Pubblicazione: (2022)
Energy Efficient Scheduling for Serverless Systems
di: Tsenos, Michail, et al.
Pubblicazione: (2024)
di: Tsenos, Michail, et al.
Pubblicazione: (2024)
KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless Computing
di: Qi, Sheng, et al.
Pubblicazione: (2026)
di: Qi, Sheng, et al.
Pubblicazione: (2026)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
di: Han, Yunhe, et al.
Pubblicazione: (2026)
di: Han, Yunhe, et al.
Pubblicazione: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
di: Zhang, Yida, et al.
Pubblicazione: (2026)
di: Zhang, Yida, et al.
Pubblicazione: (2026)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
di: Liu, Wentao, et al.
Pubblicazione: (2025)
di: Liu, Wentao, et al.
Pubblicazione: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
di: Wu, Hao, et al.
Pubblicazione: (2024)
di: Wu, Hao, et al.
Pubblicazione: (2024)
Serverless Approach to Running Resource-Intensive STAR Aligner
di: Kica, Piotr, et al.
Pubblicazione: (2025)
di: Kica, Piotr, et al.
Pubblicazione: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
di: Lin, Yanying, et al.
Pubblicazione: (2025)
di: Lin, Yanying, et al.
Pubblicazione: (2025)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
di: Carl, Natalie, et al.
Pubblicazione: (2025)
di: Carl, Natalie, et al.
Pubblicazione: (2025)
It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
di: Wu, Jing, et al.
Pubblicazione: (2025)
di: Wu, Jing, et al.
Pubblicazione: (2025)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
di: Ni, Yinan, et al.
Pubblicazione: (2025)
di: Ni, Yinan, et al.
Pubblicazione: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
di: Yu, Minchen, et al.
Pubblicazione: (2025)
di: Yu, Minchen, et al.
Pubblicazione: (2025)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
di: Yu, Minchen, et al.
Pubblicazione: (2025)
di: Yu, Minchen, et al.
Pubblicazione: (2025)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
di: Sui, Yifan, et al.
Pubblicazione: (2025)
di: Sui, Yifan, et al.
Pubblicazione: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
di: Fu, Yao, et al.
Pubblicazione: (2024)
di: Fu, Yao, et al.
Pubblicazione: (2024)
SeSeMI: Secure Serverless Model Inference on Sensitive Data
di: Hu, Guoyu, et al.
Pubblicazione: (2024)
di: Hu, Guoyu, et al.
Pubblicazione: (2024)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024)
di: He, Yiyuan, et al.
Pubblicazione: (2024)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
di: Tzenetopoulos, Achilleas, et al.
Pubblicazione: (2024)
di: Tzenetopoulos, Achilleas, et al.
Pubblicazione: (2024)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
di: O'Quinn, Austin, et al.
Pubblicazione: (2025)
di: O'Quinn, Austin, et al.
Pubblicazione: (2025)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
di: Zhao, Zhixin, et al.
Pubblicazione: (2026)
di: Zhao, Zhixin, et al.
Pubblicazione: (2026)
Energy Efficiency Support for Software Defined Networks: a Serverless Computing Approach
di: Banaie, Fatemeh, et al.
Pubblicazione: (2024)
di: Banaie, Fatemeh, et al.
Pubblicazione: (2024)
Are Unikernels Ready for Serverless on the Edge?
di: Moebius, Felix, et al.
Pubblicazione: (2024)
di: Moebius, Felix, et al.
Pubblicazione: (2024)
Caching Aided Multi-Tenant Serverless Computing
di: Qiao, Chu, et al.
Pubblicazione: (2024)
di: Qiao, Chu, et al.
Pubblicazione: (2024)
pBeeGees: A Prudent Approach to Certificate-Decoupled BFT Consensus
di: Yang, Kaiji, et al.
Pubblicazione: (2025)
di: Yang, Kaiji, et al.
Pubblicazione: (2025)
In Serverless, OS Scheduler Choice Costs Money: A Hybrid Scheduling Approach for Cheaper FaaS
di: Zhao, Yuxuan, et al.
Pubblicazione: (2024)
di: Zhao, Yuxuan, et al.
Pubblicazione: (2024)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
di: He, Yongchao, et al.
Pubblicazione: (2025)
di: He, Yongchao, et al.
Pubblicazione: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
di: Barrak, Amine, et al.
Pubblicazione: (2025)
di: Barrak, Amine, et al.
Pubblicazione: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
di: Ghosh, Himel
Pubblicazione: (2024)
di: Ghosh, Himel
Pubblicazione: (2024)
Truffle: Efficient Data Passing for Data-Intensive Serverless Workflows in the Edge-Cloud Continuum
di: Marcelino, Cynthia, et al.
Pubblicazione: (2024)
di: Marcelino, Cynthia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
di: Kong, Z. Jonny, et al.
Pubblicazione: (2025) -
Towards Resource-Efficient Serverless LLM Inference with SLINFER
di: Xu, Chuhao, et al.
Pubblicazione: (2025) -
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
di: Hui, Xinning, et al.
Pubblicazione: (2024) -
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
di: Yu, Minchen, et al.
Pubblicazione: (2023) -
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
di: Liu, Chongpeng, et al.
Pubblicazione: (2025)