TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Linfeng, Xu, Minxian, Toosi, Adel N., Ye, Kejiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
von: Jia, Haojie, et al.
Veröffentlicht: (2025)
von: Jia, Haojie, et al.
Veröffentlicht: (2025)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
von: Wu, Jingfeng, et al.
Veröffentlicht: (2024)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2024)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
von: Song, Shengye, et al.
Veröffentlicht: (2025)
von: Song, Shengye, et al.
Veröffentlicht: (2025)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
von: Bai, Haoyu, et al.
Veröffentlicht: (2024)
von: Bai, Haoyu, et al.
Veröffentlicht: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
REACH: Reinforcement Learning for Adaptive Microservice Rescheduling in the Cloud-Edge Continuum
von: Bai, Xu, et al.
Veröffentlicht: (2025)
von: Bai, Xu, et al.
Veröffentlicht: (2025)
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
von: Su, Zijie, et al.
Veröffentlicht: (2026)
von: Su, Zijie, et al.
Veröffentlicht: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
von: Song, Shengye, et al.
Veröffentlicht: (2025)
von: Song, Shengye, et al.
Veröffentlicht: (2025)
A Multi-Armed Bandit-Based Participant Selection Method for Federated Recommendation Systems
von: Liu, Jintao, et al.
Veröffentlicht: (2025)
von: Liu, Jintao, et al.
Veröffentlicht: (2025)
IntentContinuum: Using LLMs to Support Intent-Based Computing Across the Compute Continuum
von: Akbari, Negin, et al.
Veröffentlicht: (2025)
von: Akbari, Negin, et al.
Veröffentlicht: (2025)
GoldFish: Serverless Actors with Short-Term Memory State for the Edge-Cloud Continuum
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2024)
von: Marcelino, Cynthia, et al.
Veröffentlicht: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
von: Bai, Xu, et al.
Veröffentlicht: (2026)
von: Bai, Xu, et al.
Veröffentlicht: (2026)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
Towards Cloud Efficiency with Large-scale Workload Characterization
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
Ksurf: Attention Kalman Filter and Principal Component Analysis for Prediction under Highly Variable Cloud Workloads
von: Dang'ana, Michael, et al.
Veröffentlicht: (2024)
von: Dang'ana, Michael, et al.
Veröffentlicht: (2024)
Orchestrating Mixed-Criticality Cloud Workloads in Reconfigurable Manufacturing Systems
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
von: Bai, Haoyu, et al.
Veröffentlicht: (2026)
von: Bai, Haoyu, et al.
Veröffentlicht: (2026)
DistributedEstimator: Distributed Training of Quantum Neural Networks via Circuit Cutting
von: Singh, Prabhjot, et al.
Veröffentlicht: (2026)
von: Singh, Prabhjot, et al.
Veröffentlicht: (2026)
GraphFlash: Enabling Fast and Elastic Graph Processing on Serverless Infrastructure
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
Spatio-Temporal Shifting to Reduce Carbon, Water, and Land-Use Footprints of Cloud Workloads
von: Attenni, Giulio, et al.
Veröffentlicht: (2025)
von: Attenni, Giulio, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
von: Tang, Lujie, et al.
Veröffentlicht: (2024) -
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
von: Li, Xiang, et al.
Veröffentlicht: (2024) -
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
von: Hu, Kan, et al.
Veröffentlicht: (2024) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025) -
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
von: Jia, Haojie, et al.
Veröffentlicht: (2025)