A Deep Dive into the Google Cluster Workload Traces: Analyzing the Application Failure Characteristics and User Behaviors
Fuente:
arXiv
Saved in:
| Main Authors: | Bappy, Faisal Haque, Islam, Tariqul, Zaman, Tarannum Shaila, Hasan, Raiful, Caicedo, Carlos |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FASTEN: Towards a FAult-tolerant and STorage EfficieNt Cloud: Balancing Between Replication and Deduplication
by: Ahmed, Sabbir, et al.
Published: (2023)
by: Ahmed, Sabbir, et al.
Published: (2023)
SEAM: A Secure Automated and Maintainable Smart Contract Upgrade Framework
by: Hossain, Tahrim, et al.
Published: (2024)
by: Hossain, Tahrim, et al.
Published: (2024)
ConChain: A Scheme for Contention-free and Attack Resilient BlockChain
by: Bappy, Faisal Haque, et al.
Published: (2023)
by: Bappy, Faisal Haque, et al.
Published: (2023)
Maximizing Blockchain Performance: Mitigating Conflicting Transactions through Parallelism and Dependency Management
by: Bappy, Faisal Haque, et al.
Published: (2024)
by: Bappy, Faisal Haque, et al.
Published: (2024)
MRL-PoS: A Multi-agent Reinforcement Learning based Proof of Stake Consensus Algorithm for Blockchain
by: Islam, Tariqul, et al.
Published: (2023)
by: Islam, Tariqul, et al.
Published: (2023)
SWORD: A Secure LoW-Latency Offline-First Authentication and Data Sharing Scheme for Resource Constrained Distributed Networks
by: Bappy, Faisal Haque, et al.
Published: (2026)
by: Bappy, Faisal Haque, et al.
Published: (2026)
Impact of Conflicting Transactions in Blockchain: Detecting and Mitigating Potential Attacks
by: Bappy, Faisal Haque, et al.
Published: (2024)
by: Bappy, Faisal Haque, et al.
Published: (2024)
Securing Proof of Stake Blockchains: Leveraging Multi-Agent Reinforcement Learning for Detecting and Mitigating Malicious Nodes
by: Bappy, Faisal Haque, et al.
Published: (2024)
by: Bappy, Faisal Haque, et al.
Published: (2024)
Collaborative Proof-of-Work: A Secure Dynamic Approach to Fair and Efficient Blockchain Mining
by: Haque, Rizwanul, et al.
Published: (2024)
by: Haque, Rizwanul, et al.
Published: (2024)
Agentic AI Workload Characteristics
by: Yuan, Yichao, et al.
Published: (2026)
by: Yuan, Yichao, et al.
Published: (2026)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
SmartShift: A Secure and Efficient Approach to Smart Contract Migration
by: Hossain, Tahrim, et al.
Published: (2025)
by: Hossain, Tahrim, et al.
Published: (2025)
Profiling and Modeling of Power Characteristics of Leadership-Scale HPC System Workloads
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
by: Iserte, Sergio, et al.
Published: (2025)
by: Iserte, Sergio, et al.
Published: (2025)
Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
by: Zojer, Patrick, et al.
Published: (2026)
by: Zojer, Patrick, et al.
Published: (2026)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
by: Jain, Rutwik, et al.
Published: (2024)
by: Jain, Rutwik, et al.
Published: (2024)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
by: Wu, Xin, et al.
Published: (2026)
by: Wu, Xin, et al.
Published: (2026)
Dynamic Client Clustering, Bandwidth Allocation, and Workload Optimization for Semi-synchronous Federated Learning
by: Yu, Liangkun, et al.
Published: (2024)
by: Yu, Liangkun, et al.
Published: (2024)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
by: Ovi, Md Sultanul Islam
Published: (2025)
by: Ovi, Md Sultanul Islam
Published: (2025)
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
by: Ramicetty, P., et al.
Published: (2026)
by: Ramicetty, P., et al.
Published: (2026)
QONNECT: A QoS-Aware Orchestration System for Distributed Kubernetes Clusters
by: Aslan, Haci Ismail, et al.
Published: (2025)
by: Aslan, Haci Ismail, et al.
Published: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
AI Surrogate Model for Distributed Computing Workloads
by: Park, David K., et al.
Published: (2024)
by: Park, David K., et al.
Published: (2024)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Union: An Automatic Workload Manager for Accelerating Network Simulation
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Crossword: Adaptive Consensus for Dynamic Data-Heavy Workloads
by: Hu, Guanzhou, et al.
Published: (2025)
by: Hu, Guanzhou, et al.
Published: (2025)
Distributed Load Balancing with Workload-Dependent Service Rates
by: Zhang, Wenxin, et al.
Published: (2024)
by: Zhang, Wenxin, et al.
Published: (2024)
Eventually-Consistent Federated Scheduling for Data Center Workloads
by: Thiyyakat, Meghana, et al.
Published: (2023)
by: Thiyyakat, Meghana, et al.
Published: (2023)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Data Management System Analysis for Distributed Computing Workloads
by: Hsu, Kuan-Chieh, et al.
Published: (2025)
by: Hsu, Kuan-Chieh, et al.
Published: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
by: Schieffer, Gabin, et al.
Published: (2025)
by: Schieffer, Gabin, et al.
Published: (2025)
Evaluating Container Orchestration for Neuromorphic Workloads in Virtual Edge Environments
by: Pham, Huyen, et al.
Published: (2026)
by: Pham, Huyen, et al.
Published: (2026)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
by: Zhu, Botao, et al.
Published: (2025)
by: Zhu, Botao, et al.
Published: (2025)
Similar Items
-
FASTEN: Towards a FAult-tolerant and STorage EfficieNt Cloud: Balancing Between Replication and Deduplication
by: Ahmed, Sabbir, et al.
Published: (2023) -
SEAM: A Secure Automated and Maintainable Smart Contract Upgrade Framework
by: Hossain, Tahrim, et al.
Published: (2024) -
ConChain: A Scheme for Contention-free and Attack Resilient BlockChain
by: Bappy, Faisal Haque, et al.
Published: (2023) -
Maximizing Blockchain Performance: Mitigating Conflicting Transactions through Parallelism and Dependency Management
by: Bappy, Faisal Haque, et al.
Published: (2024) -
MRL-PoS: A Multi-agent Reinforcement Learning based Proof of Stake Consensus Algorithm for Blockchain
by: Islam, Tariqul, et al.
Published: (2023)