Saved in:
| Main Authors: | Wu, Jing, Wang, Lin, Deng, Quanfeng, Yu, Chen, Zhang, Dong, Yan, Bingheng, Liu, Fangming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.14320 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
by: Lv, Cunchi, et al.
Published: (2025)
by: Lv, Cunchi, et al.
Published: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
by: Pagonas, Nikos, et al.
Published: (2025)
by: Pagonas, Nikos, et al.
Published: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023)
by: Yu, Minchen, et al.
Published: (2023)
AARC: Automated Affinity-aware Resource Configuration for Serverless Workflows
by: Jin, Lingxiao, et al.
Published: (2025)
by: Jin, Lingxiao, et al.
Published: (2025)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
GeoFF: Federated Serverless Workflows with Data Pre-Fetching
by: Carl, Natalie, et al.
Published: (2024)
by: Carl, Natalie, et al.
Published: (2024)
sAirflow: Adopting Serverless in a Legacy Workflow Scheduler
by: Mikina, Filip, et al.
Published: (2024)
by: Mikina, Filip, et al.
Published: (2024)
Jiagu: Optimizing Serverless Computing Resource Utilization with Harmonized Efficiency and Practicability
by: Liu, Qingyuan, et al.
Published: (2024)
by: Liu, Qingyuan, et al.
Published: (2024)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
by: Tzenetopoulos, Achilleas, et al.
Published: (2024)
by: Tzenetopoulos, Achilleas, et al.
Published: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
by: Lin, Yanying, et al.
Published: (2025)
by: Lin, Yanying, et al.
Published: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Cosmos: A Cost Model for Serverless Workflows in the 3D Compute Continuum
by: Marcelino, Cynthia, et al.
Published: (2025)
by: Marcelino, Cynthia, et al.
Published: (2025)
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
by: Farahani, Reza, et al.
Published: (2026)
by: Farahani, Reza, et al.
Published: (2026)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)
by: Peng, You, et al.
Published: (2026)
OTAS: An Elastic Transformer Serving System via Token Adaptation
by: Chen, Jinyu, et al.
Published: (2024)
by: Chen, Jinyu, et al.
Published: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
by: Hui, Xinning, et al.
Published: (2024)
by: Hui, Xinning, et al.
Published: (2024)
Databelt: A Continuous Data Path for Serverless Workflows in the 3D Compute Continuum
by: Marcelino, Cynthia, et al.
Published: (2025)
by: Marcelino, Cynthia, et al.
Published: (2025)
Truffle: Efficient Data Passing for Data-Intensive Serverless Workflows in the Edge-Cloud Continuum
by: Marcelino, Cynthia, et al.
Published: (2024)
by: Marcelino, Cynthia, et al.
Published: (2024)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
by: Dai, Yinwei, et al.
Published: (2025)
by: Dai, Yinwei, et al.
Published: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
by: Du, Boxiao, et al.
Published: (2026)
by: Du, Boxiao, et al.
Published: (2026)
Dependency-aware Resource Allocation for Serverless Functions at the Edge
by: Baresi, Luciano, et al.
Published: (2023)
by: Baresi, Luciano, et al.
Published: (2023)
Serverless Approach to Running Resource-Intensive STAR Aligner
by: Kica, Piotr, et al.
Published: (2025)
by: Kica, Piotr, et al.
Published: (2025)
Serverless Everywhere: A Comparative Analysis of WebAssembly Workflows Across Browser, Edge, and Cloud
by: Colosi, Mario, et al.
Published: (2025)
by: Colosi, Mario, et al.
Published: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
by: Lin, Zejia, et al.
Published: (2025)
by: Lin, Zejia, et al.
Published: (2025)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
ICPS: Real-Time Resource Configuration for Cloud Serverless Functions Considering Affinity
by: Chen, Long, et al.
Published: (2025)
by: Chen, Long, et al.
Published: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
by: Zhang, Yuning, et al.
Published: (2026)
by: Zhang, Yuning, et al.
Published: (2026)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
by: Schmid, Larissa, et al.
Published: (2024)
by: Schmid, Larissa, et al.
Published: (2024)
Adaptive Resource Allocation for Workflow Containerization on Kubernetes
by: Shan, Chenggang, et al.
Published: (2023)
by: Shan, Chenggang, et al.
Published: (2023)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
by: Wu, Z., et al.
Published: (2025)
by: Wu, Z., et al.
Published: (2025)
Zenix: Efficient Execution of Bulky Serverless Applications
by: Guo, Zhiyuan, et al.
Published: (2022)
by: Guo, Zhiyuan, et al.
Published: (2022)
MoEless: Efficient MoE LLM Serving via Serverless Computing
by: Yu, Hanfei, et al.
Published: (2026)
by: Yu, Hanfei, et al.
Published: (2026)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
by: Shen, Haiying, et al.
Published: (2024)
by: Shen, Haiying, et al.
Published: (2024)
Similar Items
-
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025) -
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
by: Lv, Cunchi, et al.
Published: (2025) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024) -
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025) -
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
by: Pagonas, Nikos, et al.
Published: (2025)