Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling
Fuente:
arXiv
Guardado en:
| Autores principales: | Jin, Yihong, Yang, Ze |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
por: Zehra, Sehar, et al.
Publicado: (2025)
por: Zehra, Sehar, et al.
Publicado: (2025)
AAFLOW: Scalable Patterns for Agentic AI Workflows
por: Sarker, Arup Kumar, et al.
Publicado: (2026)
por: Sarker, Arup Kumar, et al.
Publicado: (2026)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
por: Will, Jonathan, et al.
Publicado: (2025)
por: Will, Jonathan, et al.
Publicado: (2025)
Scalable overset computation between a forest-of-octrees- and an arbitrary distributed parallel mesh
por: Brandt, Hannes, et al.
Publicado: (2026)
por: Brandt, Hannes, et al.
Publicado: (2026)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
por: Sun, He, et al.
Publicado: (2026)
por: Sun, He, et al.
Publicado: (2026)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
por: Nguyen, Ngoc Hung, et al.
Publicado: (2025)
por: Nguyen, Ngoc Hung, et al.
Publicado: (2025)
Sublinear-Time Sampling of Spanning Trees in the Congested Clique
por: Pemmaraju, Sriram V., et al.
Publicado: (2024)
por: Pemmaraju, Sriram V., et al.
Publicado: (2024)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
por: Gao, Yuxuan, et al.
Publicado: (2026)
por: Gao, Yuxuan, et al.
Publicado: (2026)
A Note on Solving Problems of Substantially Super-linear Complexity in $N^{o(1)}$ Rounds of the Congested Clique
por: Lingas, Andrzej
Publicado: (2024)
por: Lingas, Andrzej
Publicado: (2024)
Efficient Construction of Large Search Spaces for Auto-Tuning
por: Willemsen, Floris-Jan, et al.
Publicado: (2025)
por: Willemsen, Floris-Jan, et al.
Publicado: (2025)
Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging
por: Wu, Zihan, et al.
Publicado: (2024)
por: Wu, Zihan, et al.
Publicado: (2024)
Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains
por: Eduardo, González Trigueros Jesús, et al.
Publicado: (2025)
por: Eduardo, González Trigueros Jesús, et al.
Publicado: (2025)
Deterministic Fault-Tolerant Local Load Balancing and its Applications against Adaptive Adversaries
por: Kowalski, Dariusz R., et al.
Publicado: (2025)
por: Kowalski, Dariusz R., et al.
Publicado: (2025)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
por: Mohammad, Noor Islam S.
Publicado: (2025)
por: Mohammad, Noor Islam S.
Publicado: (2025)
Near-Optimal Wafer-Scale Reduce
por: Luczynski, Piotr, et al.
Publicado: (2024)
por: Luczynski, Piotr, et al.
Publicado: (2024)
Obfuscated Consensus
por: Aspnes, James, et al.
Publicado: (2025)
por: Aspnes, James, et al.
Publicado: (2025)
Why Canonical Rounds Fail for Optimal Byzantine Resilience
por: Attiya, Hagit, et al.
Publicado: (2025)
por: Attiya, Hagit, et al.
Publicado: (2025)
Improving Efficiency in Near-State and State-Optimal Self-Stabilising Leader Election Population Protocols
por: Gąsieniec, Leszek, et al.
Publicado: (2025)
por: Gąsieniec, Leszek, et al.
Publicado: (2025)
Anonymous Self-Stabilising Localisation via Spatial Population Protocols
por: Gąsieniec, Leszek, et al.
Publicado: (2024)
por: Gąsieniec, Leszek, et al.
Publicado: (2024)
An Analysis of Avalanche Consensus
por: Amores-Sesar, Ignacio, et al.
Publicado: (2024)
por: Amores-Sesar, Ignacio, et al.
Publicado: (2024)
The consensus number of a shift register equals its width
por: Aspnes, James
Publicado: (2025)
por: Aspnes, James
Publicado: (2025)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
Clock Synchronization Is Almost Impossible with Bounded Memory
por: Charron-Bost, Bernadette, et al.
Publicado: (2024)
por: Charron-Bost, Bernadette, et al.
Publicado: (2024)
RadiK: Scalable and Optimized GPU-Parallel Radix Top-K Selection
por: Li, Yifei, et al.
Publicado: (2025)
por: Li, Yifei, et al.
Publicado: (2025)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
por: Putra, Jody Almaida
Publicado: (2026)
por: Putra, Jody Almaida
Publicado: (2026)
Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
por: Chundru, Jagadeesh
Publicado: (2026)
por: Chundru, Jagadeesh
Publicado: (2026)
Optimizing edge AI models on HPC systems with the edge in the loop
por: Aach, Marcel, et al.
Publicado: (2025)
por: Aach, Marcel, et al.
Publicado: (2025)
Low-Bandwidth Matrix Multiplication: Faster Algorithms and More General Forms of Sparsity
por: Gupta, Chetan, et al.
Publicado: (2024)
por: Gupta, Chetan, et al.
Publicado: (2024)
Augmenting the FedProx Algorithm by Minimizing Convergence
por: Sarkar, Anomitra, et al.
Publicado: (2024)
por: Sarkar, Anomitra, et al.
Publicado: (2024)
N2N: A Parallel Framework for Large-Scale MILP under Distributed Memory
por: Wang, Longfei, et al.
Publicado: (2025)
por: Wang, Longfei, et al.
Publicado: (2025)
Boolean Matrix Multiplication for Highly Clustered Data on the Congested Clique
por: Lingas, Andrzej
Publicado: (2024)
por: Lingas, Andrzej
Publicado: (2024)
Learning Interpretable Scheduling Algorithms for Data Processing Clusters
por: Hu, Zhibo, et al.
Publicado: (2024)
por: Hu, Zhibo, et al.
Publicado: (2024)
Scalable Mesh Coupling for Atmospheric Wave Simulation
por: Brandt, Hannes, et al.
Publicado: (2026)
por: Brandt, Hannes, et al.
Publicado: (2026)
Heuristic Search Space Partitioning for Low-Latency Multi-Tenant Cloud Queries
por: Pathak, Prashant Kumar, et al.
Publicado: (2026)
por: Pathak, Prashant Kumar, et al.
Publicado: (2026)
When Does Global Attention Help? A Unified Empirical Study on Atomistic Graph Learning
por: Chowdhury, Arindam, et al.
Publicado: (2025)
por: Chowdhury, Arindam, et al.
Publicado: (2025)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
por: Woisetschläger, Herbert, et al.
Publicado: (2023)
por: Woisetschläger, Herbert, et al.
Publicado: (2023)
Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon
por: Khan, Shamsher
Publicado: (2026)
por: Khan, Shamsher
Publicado: (2026)
Distributed Rhombus Formation of Sliding Squares
por: Kostitsyna, Irina, et al.
Publicado: (2025)
por: Kostitsyna, Irina, et al.
Publicado: (2025)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
por: Dutta, Yuvraj, et al.
Publicado: (2025)
por: Dutta, Yuvraj, et al.
Publicado: (2025)
Ejemplares similares
-
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
por: Sharma, Aasish Kumar, et al.
Publicado: (2025) -
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
por: Zehra, Sehar, et al.
Publicado: (2025) -
AAFLOW: Scalable Patterns for Agentic AI Workflows
por: Sarker, Arup Kumar, et al.
Publicado: (2026) -
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
por: Will, Jonathan, et al.
Publicado: (2025) -
Scalable overset computation between a forest-of-octrees- and an arbitrary distributed parallel mesh
por: Brandt, Hannes, et al.
Publicado: (2026)