Gespeichert in:
| Hauptverfasser: | Wu, Jing, Wang, Lin, Deng, Quanfeng, Yu, Chen, Zhang, Dong, Yan, Bingheng, Liu, Fangming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.14320 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
AARC: Automated Affinity-aware Resource Configuration for Serverless Workflows
von: Jin, Lingxiao, et al.
Veröffentlicht: (2025)
von: Jin, Lingxiao, et al.
Veröffentlicht: (2025)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
GeoFF: Federated Serverless Workflows with Data Pre-Fetching
von: Carl, Natalie, et al.
Veröffentlicht: (2024)
von: Carl, Natalie, et al.
Veröffentlicht: (2024)
sAirflow: Adopting Serverless in a Legacy Workflow Scheduler
von: Mikina, Filip, et al.
Veröffentlicht: (2024)
von: Mikina, Filip, et al.
Veröffentlicht: (2024)
Jiagu: Optimizing Serverless Computing Resource Utilization with Harmonized Efficiency and Practicability
von: Liu, Qingyuan, et al.
Veröffentlicht: (2024)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2024)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
von: Tzenetopoulos, Achilleas, et al.
Veröffentlicht: (2024)
von: Tzenetopoulos, Achilleas, et al.
Veröffentlicht: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Cosmos: A Cost Model for Serverless Workflows in the 3D Compute Continuum
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
von: Farahani, Reza, et al.
Veröffentlicht: (2026)
von: Farahani, Reza, et al.
Veröffentlicht: (2026)
Software Resource Disaggregation for HPC with Serverless Computing
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
OTAS: An Elastic Transformer Serving System via Token Adaptation
von: Chen, Jinyu, et al.
Veröffentlicht: (2024)
von: Chen, Jinyu, et al.
Veröffentlicht: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
Databelt: A Continuous Data Path for Serverless Workflows in the 3D Compute Continuum
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2025)
Truffle: Efficient Data Passing for Data-Intensive Serverless Workflows in the Edge-Cloud Continuum
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2024)
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2024)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
Dependency-aware Resource Allocation for Serverless Functions at the Edge
von: Baresi, Luciano, et al.
Veröffentlicht: (2023)
von: Baresi, Luciano, et al.
Veröffentlicht: (2023)
Serverless Approach to Running Resource-Intensive STAR Aligner
von: Kica, Piotr, et al.
Veröffentlicht: (2025)
von: Kica, Piotr, et al.
Veröffentlicht: (2025)
Serverless Everywhere: A Comparative Analysis of WebAssembly Workflows Across Browser, Edge, and Cloud
von: Colosi, Mario, et al.
Veröffentlicht: (2025)
von: Colosi, Mario, et al.
Veröffentlicht: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
ICPS: Real-Time Resource Configuration for Cloud Serverless Functions Considering Affinity
von: Chen, Long, et al.
Veröffentlicht: (2025)
von: Chen, Long, et al.
Veröffentlicht: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
von: Schmid, Larissa, et al.
Veröffentlicht: (2024)
von: Schmid, Larissa, et al.
Veröffentlicht: (2024)
Adaptive Resource Allocation for Workflow Containerization on Kubernetes
von: Shan, Chenggang, et al.
Veröffentlicht: (2023)
von: Shan, Chenggang, et al.
Veröffentlicht: (2023)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
von: Wu, Z., et al.
Veröffentlicht: (2025)
von: Wu, Z., et al.
Veröffentlicht: (2025)
Zenix: Efficient Execution of Bulky Serverless Applications
von: Guo, Zhiyuan, et al.
Veröffentlicht: (2022)
von: Guo, Zhiyuan, et al.
Veröffentlicht: (2022)
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025) -
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
von: Lv, Cunchi, et al.
Veröffentlicht: (2025) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024) -
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025) -
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)