FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Minghe, Schirmer, Trever, Malekabbasi, Mohammadreza, Bermbach, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GeoFaaS: An Edge-to-Cloud FaaS Platform
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2024)
DisCEdge: Distributed Context Management for Large Language Models at the Edge
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
LLM4FaaS: No-Code Application Development using LLMs and FaaS
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
Towards a Testbed for Scalable FaaS Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2025)
von: Schirmer, Trever, et al.
Veröffentlicht: (2025)
ProFaaStinate: Delaying Serverless Function Calls to Optimize Platform Performance
von: Schirmer, Trever, et al.
Veröffentlicht: (2023)
von: Schirmer, Trever, et al.
Veröffentlicht: (2023)
ElastiBench: Scalable Continuous Benchmarking on Cloud FaaS Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2024)
von: Schirmer, Trever, et al.
Veröffentlicht: (2024)
Konflux: Optimized Function Fusion for Serverless Applications
von: Kowallik, Niklas, et al.
Veröffentlicht: (2026)
von: Kowallik, Niklas, et al.
Veröffentlicht: (2026)
Umbilical Choir: Automated Live Testing for Edge-To-Cloud FaaS Applications
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
Trabant: A Serverless Architecture for Multi-Tenant Orbital Edge Computing
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2025)
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2025)
GeoFF: Federated Serverless Workflows with Data Pre-Fetching
von: Carl, Natalie, et al.
Veröffentlicht: (2024)
von: Carl, Natalie, et al.
Veröffentlicht: (2024)
Fusionize++: Improving Serverless Application Performance Using Dynamic Task Inlining and Infrastructure Optimization
von: Schirmer, Trever, et al.
Veröffentlicht: (2023)
von: Schirmer, Trever, et al.
Veröffentlicht: (2023)
Exploring Influence Factors on LLM Suitability for No-Code Development of End User IoT Applications
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
Multi-Event Triggers for Serverless Computing
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
Serverless Abstractions for Short-Running, Lightweight Streams
von: Carl, Natalie, et al.
Veröffentlicht: (2026)
von: Carl, Natalie, et al.
Veröffentlicht: (2026)
New Kids: An Architecture and Performance Investigation of Second-Generation Serverless Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2026)
von: Schirmer, Trever, et al.
Veröffentlicht: (2026)
Application-Centric Benchmarking of Distributed FaaS Platforms using BeFaaS
von: Grambow, Martin, et al.
Veröffentlicht: (2023)
von: Grambow, Martin, et al.
Veröffentlicht: (2023)
Minos: Exploiting Cloud Performance Variation with Function-as-a-Service Instance Selection
von: Schirmer, Trever, et al.
Veröffentlicht: (2025)
von: Schirmer, Trever, et al.
Veröffentlicht: (2025)
MUSE: Multi-Tenant Model Serving With Seamless Model Updates
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
Caching Aided Multi-Tenant Serverless Computing
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
Are Unikernels Ready for Serverless on the Edge?
von: Moebius, Felix, et al.
Veröffentlicht: (2024)
von: Moebius, Felix, et al.
Veröffentlicht: (2024)
Komet: A Serverless Platform for Low-Earth Orbit Edge Services
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2024)
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2024)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
FaasMeter: Energy-First Serverless Computing
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
Increasing Efficiency and Result Reliability of Continuous Benchmarking for FaaS Applications
von: Rese, Tim C., et al.
Veröffentlicht: (2024)
von: Rese, Tim C., et al.
Veröffentlicht: (2024)
Provuse: Platform-Side Function Fusion for Performance and Efficiency in FaaS Environments
von: Kowallik, Niklas, et al.
Veröffentlicht: (2026)
von: Kowallik, Niklas, et al.
Veröffentlicht: (2026)
Flight: A FaaS-Based Framework for Complex and Hierarchical Federated Learning
von: Hudson, Nathaniel, et al.
Veröffentlicht: (2024)
von: Hudson, Nathaniel, et al.
Veröffentlicht: (2024)
FaaS Is Not Enough: Serverless Handling of Burst-Parallel Jobs
von: Barcelona-Pons, Daniel, et al.
Veröffentlicht: (2024)
von: Barcelona-Pons, Daniel, et al.
Veröffentlicht: (2024)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Predicting Temporal Aspects of Movement for Predictive Replication in Fog Environments
von: Balitzki, Emil, et al.
Veröffentlicht: (2023)
von: Balitzki, Emil, et al.
Veröffentlicht: (2023)
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
SMoFi: Step-wise Momentum Fusion for Split Federated Learning on Heterogeneous Data
von: Yang, Mingkun, et al.
Veröffentlicht: (2025)
von: Yang, Mingkun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GeoFaaS: An Edge-to-Cloud FaaS Platform
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2024) -
DisCEdge: Distributed Context Management for Large Language Models at the Edge
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025) -
LLM4FaaS: No-Code Application Development using LLMs and FaaS
von: Wang, Minghe, et al.
Veröffentlicht: (2025) -
Towards a Testbed for Scalable FaaS Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2025) -
ProFaaStinate: Delaying Serverless Function Calls to Optimize Platform Performance
von: Schirmer, Trever, et al.
Veröffentlicht: (2023)