MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Jiaang, Qian, Shiyou, Yang, Dingyu, Hu, Hanwen, Cao, Jian, Xue, Guangtao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
von: Yang, Dingyu, et al.
Veröffentlicht: (2024)
von: Yang, Dingyu, et al.
Veröffentlicht: (2024)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
von: Hua, Qin, et al.
Veröffentlicht: (2024)
von: Hua, Qin, et al.
Veröffentlicht: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
von: Yang, Dingyu, et al.
Veröffentlicht: (2026)
von: Yang, Dingyu, et al.
Veröffentlicht: (2026)
Komet: A Serverless Platform for Low-Earth Orbit Edge Services
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2024)
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
GreenWhisk: Emission-Aware Computing for Serverless Platform
von: Serenari, Jayden, et al.
Veröffentlicht: (2024)
von: Serenari, Jayden, et al.
Veröffentlicht: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
A Hodge-Based Framework for Service Operational Analysis in Serverless Platforms
von: Reali, Gianluca, et al.
Veröffentlicht: (2026)
von: Reali, Gianluca, et al.
Veröffentlicht: (2026)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
von: Wu, Z., et al.
Veröffentlicht: (2025)
von: Wu, Z., et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
von: Copik, Marcin, et al.
Veröffentlicht: (2022)
von: Copik, Marcin, et al.
Veröffentlicht: (2022)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
Saarthi: An End-to-End Intelligent Platform for Optimising Distributed Serverless Workloads
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2025)
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2025)
New Kids: An Architecture and Performance Investigation of Second-Generation Serverless Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2026)
von: Schirmer, Trever, et al.
Veröffentlicht: (2026)
FaaSMT: Lightweight Serverless Framework for Intrusion Detection Using Merkle Tree and Task Inlining
von: Li, Chuang, et al.
Veröffentlicht: (2025)
von: Li, Chuang, et al.
Veröffentlicht: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
ProFaaStinate: Delaying Serverless Function Calls to Optimize Platform Performance
von: Schirmer, Trever, et al.
Veröffentlicht: (2023)
von: Schirmer, Trever, et al.
Veröffentlicht: (2023)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
von: Ni, Yinan, et al.
Veröffentlicht: (2025)
von: Ni, Yinan, et al.
Veröffentlicht: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
Tutorial: Object as a Service (OaaS) Serverless Cloud Computing Paradigm
von: Lertpongrujikorn, Pawissanutt, et al.
Veröffentlicht: (2024)
von: Lertpongrujikorn, Pawissanutt, et al.
Veröffentlicht: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
SeSeMI: Secure Serverless Model Inference on Sensitive Data
von: Hu, Guoyu, et al.
Veröffentlicht: (2024)
von: Hu, Guoyu, et al.
Veröffentlicht: (2024)
Shard the Gradient, Scale the Model: Serverless Federated Aggregation via Gradient Partitioning
von: Barrak, Amine
Veröffentlicht: (2026)
von: Barrak, Amine
Veröffentlicht: (2026)
LIFL: A Lightweight, Event-driven Serverless Platform for Federated Learning
von: Qi, Shixiong, et al.
Veröffentlicht: (2024)
von: Qi, Shixiong, et al.
Veröffentlicht: (2024)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
Cooperative Inference with Interleaved Operator Partitioning for CNNs
von: Liu, Zhibang, et al.
Veröffentlicht: (2024)
von: Liu, Zhibang, et al.
Veröffentlicht: (2024)
Resource Allocation of Industry 4.0 Micro-Service Applications across Serverless Fog Federation
von: Hussain, Razin Farhan, et al.
Veröffentlicht: (2024)
von: Hussain, Razin Farhan, et al.
Veröffentlicht: (2024)
Are Unikernels Ready for Serverless on the Edge?
von: Moebius, Felix, et al.
Veröffentlicht: (2024)
von: Moebius, Felix, et al.
Veröffentlicht: (2024)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
von: Liu, Zhibang, et al.
Veröffentlicht: (2025)
von: Liu, Zhibang, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning (DRL)-based Methods for Serverless Stream Processing Engines: A Vision, Architectural Elements, and Future Directions
von: Read, Maria R., et al.
Veröffentlicht: (2024)
von: Read, Maria R., et al.
Veröffentlicht: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
Multi-Event Triggers for Serverless Computing
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
Raptor: Distributed Scheduling for Serverless Functions
von: Exton, Kevin, et al.
Veröffentlicht: (2024)
von: Exton, Kevin, et al.
Veröffentlicht: (2024)
Energy Efficient Scheduling for Serverless Systems
von: Tsenos, Michail, et al.
Veröffentlicht: (2024)
von: Tsenos, Michail, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
von: Yang, Dingyu, et al.
Veröffentlicht: (2024) -
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
von: Hua, Qin, et al.
Veröffentlicht: (2024) -
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
von: Duan, Jiaang, et al.
Veröffentlicht: (2025) -
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
von: Yang, Dingyu, et al.
Veröffentlicht: (2026) -
Komet: A Serverless Platform for Low-Earth Orbit Edge Services
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2024)