Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Barrak, Amine, Ksontini, Emna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Shard the Gradient, Scale the Model: Serverless Federated Aggregation via Gradient Partitioning
von: Barrak, Amine
Veröffentlicht: (2026)
von: Barrak, Amine
Veröffentlicht: (2026)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
Shabari: Delayed Decision-Making for Faster and Efficient Serverless Functions
von: Sinha, Prasoon, et al.
Veröffentlicht: (2024)
von: Sinha, Prasoon, et al.
Veröffentlicht: (2024)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
Reinforcement Learning-Based Dynamic Management of Structured Parallel Farm Skeletons on Serverless Platforms
von: Li, Lanpei, et al.
Veröffentlicht: (2026)
von: Li, Lanpei, et al.
Veröffentlicht: (2026)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
von: Ning, Rui, et al.
Veröffentlicht: (2026)
von: Ning, Rui, et al.
Veröffentlicht: (2026)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
von: Zhang, Zongshun, et al.
Veröffentlicht: (2025)
von: Zhang, Zongshun, et al.
Veröffentlicht: (2025)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
Apodotiko: Enabling Efficient Serverless Federated Learning in Heterogeneous Environments
von: Chadha, Mohak, et al.
Veröffentlicht: (2024)
von: Chadha, Mohak, et al.
Veröffentlicht: (2024)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
von: Prabhakar, Rohan Baskar, et al.
Veröffentlicht: (2024)
von: Prabhakar, Rohan Baskar, et al.
Veröffentlicht: (2024)
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
von: Rajbhandari, Samyam, et al.
Veröffentlicht: (2025)
von: Rajbhandari, Samyam, et al.
Veröffentlicht: (2025)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
von: Su, Jianchang, et al.
Veröffentlicht: (2026)
von: Su, Jianchang, et al.
Veröffentlicht: (2026)
DynLP: Parallel Dynamic Batch Update for Label Propagation in Semi-Supervised Learning
von: Shovan, S M, et al.
Veröffentlicht: (2026)
von: Shovan, S M, et al.
Veröffentlicht: (2026)
Context Parallelism for Scalable Million-Token Inference
von: Yang, Amy, et al.
Veröffentlicht: (2024)
von: Yang, Amy, et al.
Veröffentlicht: (2024)
High-Performance Serverless Computing: A Systematic Literature Review on Serverless for HPC, AI, and Big Data
von: Besozzi, Valerio, et al.
Veröffentlicht: (2026)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2026)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
von: Khare, Alind, et al.
Veröffentlicht: (2023)
von: Khare, Alind, et al.
Veröffentlicht: (2023)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
von: Luan, Frank Sifei, et al.
Veröffentlicht: (2025)
von: Luan, Frank Sifei, et al.
Veröffentlicht: (2025)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
von: Tian, Jian, et al.
Veröffentlicht: (2025)
von: Tian, Jian, et al.
Veröffentlicht: (2025)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
von: Parikh, Dhruv, et al.
Veröffentlicht: (2025)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
LIFL: A Lightweight, Event-driven Serverless Platform for Federated Learning
von: Qi, Shixiong, et al.
Veröffentlicht: (2024)
von: Qi, Shixiong, et al.
Veröffentlicht: (2024)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
von: Wang, Chong, et al.
Veröffentlicht: (2026)
von: Wang, Chong, et al.
Veröffentlicht: (2026)
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
von: Bai, Huawei, et al.
Veröffentlicht: (2025)
von: Bai, Huawei, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
tf.data service: A Case for Disaggregating ML Input Data Processing
von: Audibert, Andrew, et al.
Veröffentlicht: (2022)
von: Audibert, Andrew, et al.
Veröffentlicht: (2022)
Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
von: Yoon, Daegun, et al.
Veröffentlicht: (2024)
von: Yoon, Daegun, et al.
Veröffentlicht: (2024)
PaSE: Parallelization Strategies for Efficient DNN Training
von: Elango, Venmugil
Veröffentlicht: (2024)
von: Elango, Venmugil
Veröffentlicht: (2024)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
von: Niu, Yifan, et al.
Veröffentlicht: (2026)
von: Niu, Yifan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Shard the Gradient, Scale the Model: Serverless Federated Aggregation via Gradient Partitioning
von: Barrak, Amine
Veröffentlicht: (2026) -
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
von: Barrak, Amine, et al.
Veröffentlicht: (2025) -
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
von: Sui, Yifan, et al.
Veröffentlicht: (2025) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024) -
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024)