CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shahbazinia, Amirhossein, Huang, Darong, Costero, Luis, Atienza, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?
von: Dou, Xinglei, et al.
Veröffentlicht: (2025)
von: Dou, Xinglei, et al.
Veröffentlicht: (2025)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
von: Costello, Ian J., et al.
Veröffentlicht: (2020)
von: Costello, Ian J., et al.
Veröffentlicht: (2020)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
von: Dutta, Akash, et al.
Veröffentlicht: (2024)
von: Dutta, Akash, et al.
Veröffentlicht: (2024)
Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
The SAP Cloud Infrastructure Dataset: A Reality Check of Scheduling and Placement of VMs in Cloud Computing
von: Uhlig, Arno, et al.
Veröffentlicht: (2025)
von: Uhlig, Arno, et al.
Veröffentlicht: (2025)
Cloud Resource Allocation with Convex Optimization
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
Usability Evaluation of Cloud for HPC Applications
von: Sochat, Vanessa, et al.
Veröffentlicht: (2025)
von: Sochat, Vanessa, et al.
Veröffentlicht: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
von: Ding, Shiwei, et al.
Veröffentlicht: (2025)
von: Ding, Shiwei, et al.
Veröffentlicht: (2025)
Bridding OT and PaaS in Edge-to-Cloud Continuum
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
Optimal Configuration of API Resources in Cloud Native Computing
von: Truyen, Eddy, et al.
Veröffentlicht: (2025)
von: Truyen, Eddy, et al.
Veröffentlicht: (2025)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
LLMPerf: GPU Performance Modeling meets Large Language Models
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
Sampling in Cloud Benchmarking: A Critical Review and Methodological Guidelines
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
von: Oztop, Beste, et al.
Veröffentlicht: (2026)
von: Oztop, Beste, et al.
Veröffentlicht: (2026)
On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems
von: Loghin, Dumitrel, et al.
Veröffentlicht: (2025)
von: Loghin, Dumitrel, et al.
Veröffentlicht: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
von: Huang, Zhaolan, et al.
Veröffentlicht: (2025)
von: Huang, Zhaolan, et al.
Veröffentlicht: (2025)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
von: John, Chelsea Maria, et al.
Veröffentlicht: (2024)
von: John, Chelsea Maria, et al.
Veröffentlicht: (2024)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
MegaFold: System-Level Optimizations for Accelerating Protein Structure Prediction Models
von: La, Hoa, et al.
Veröffentlicht: (2025)
von: La, Hoa, et al.
Veröffentlicht: (2025)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
von: Wu, Dan, et al.
Veröffentlicht: (2023)
von: Wu, Dan, et al.
Veröffentlicht: (2023)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?
von: Dou, Xinglei, et al.
Veröffentlicht: (2025) -
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
von: Xu, Guanyu, et al.
Veröffentlicht: (2025) -
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025) -
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024) -
Analytics of Longitudinal System Monitoring Data for Performance Prediction
von: Costello, Ian J., et al.
Veröffentlicht: (2020)