Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zarkadas, Ioannis, Tomlinson, Amanda, Cidon, Asaf, Kasikci, Baris, Weisse, Ofir |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
por: Apanasevich, L., et al.
Publicado: (2024)
por: Apanasevich, L., et al.
Publicado: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
por: Bogart, Christopher, et al.
Publicado: (2025)
por: Bogart, Christopher, et al.
Publicado: (2025)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
por: Debnath, Shimul, et al.
Publicado: (2026)
por: Debnath, Shimul, et al.
Publicado: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
por: Jeffery, Andrew, et al.
Publicado: (2024)
por: Jeffery, Andrew, et al.
Publicado: (2024)
LLload: Simplifying Real-Time Job Monitoring for HPC Users
por: Byun, Chansup, et al.
Publicado: (2024)
por: Byun, Chansup, et al.
Publicado: (2024)
PICO: Performance Insights for Collective Operations
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
por: Vatsavai, Sairam Sri, et al.
Publicado: (2025)
por: Vatsavai, Sairam Sri, et al.
Publicado: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
por: McDonald, Jesse, et al.
Publicado: (2024)
por: McDonald, Jesse, et al.
Publicado: (2024)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
por: Wang, Yidi, et al.
Publicado: (2024)
por: Wang, Yidi, et al.
Publicado: (2024)
A Performance Analysis of BFT Consensus for Blockchains
por: Chan, J. D., et al.
Publicado: (2024)
por: Chan, J. D., et al.
Publicado: (2024)
Scalable GPU Performance Variability Analysis framework
por: Lahiry, Ankur, et al.
Publicado: (2025)
por: Lahiry, Ankur, et al.
Publicado: (2025)
Automated Programmatic Performance Analysis of Parallel Programs
por: Cankur, Onur, et al.
Publicado: (2024)
por: Cankur, Onur, et al.
Publicado: (2024)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
por: Wang, Yuxin, et al.
Publicado: (2024)
por: Wang, Yuxin, et al.
Publicado: (2024)
Temporal Load Imbalance on Ondes3D Seismic Simulator for Different Multicore Architectures
por: Solórzano, Ana Luisa Veroneze, et al.
Publicado: (2024)
por: Solórzano, Ana Luisa Veroneze, et al.
Publicado: (2024)
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
por: Henning, Sören, et al.
Publicado: (2025)
por: Henning, Sören, et al.
Publicado: (2025)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
por: Dutilleul, Alban, et al.
Publicado: (2024)
por: Dutilleul, Alban, et al.
Publicado: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
por: de Morais, Gustavo, et al.
Publicado: (2025)
por: de Morais, Gustavo, et al.
Publicado: (2025)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
por: Aqasizade, Hossein, et al.
Publicado: (2024)
por: Aqasizade, Hossein, et al.
Publicado: (2024)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
por: Banegas-Luna, Antonio Jesús, et al.
Publicado: (2025)
por: Banegas-Luna, Antonio Jesús, et al.
Publicado: (2025)
Taking GPU Programming Models to Task for Performance Portability
por: Davis, Joshua H., et al.
Publicado: (2024)
por: Davis, Joshua H., et al.
Publicado: (2024)
eBPF-Based Instrumentation for Generalisable Diagnosis of Performance Degradation
por: Landau, Diogo, et al.
Publicado: (2025)
por: Landau, Diogo, et al.
Publicado: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026)
por: Davis, Joshua H., et al.
Publicado: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
por: Pilliat, Emmanuel
Publicado: (2026)
por: Pilliat, Emmanuel
Publicado: (2026)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
por: Arif, Moiz, et al.
Publicado: (2026)
por: Arif, Moiz, et al.
Publicado: (2026)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
por: Afzal, Ayesha, et al.
Publicado: (2025)
por: Afzal, Ayesha, et al.
Publicado: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
por: Ramesh, Risshab Srinivas
Publicado: (2024)
por: Ramesh, Risshab Srinivas
Publicado: (2024)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
por: Grbic, Dragana
Publicado: (2026)
por: Grbic, Dragana
Publicado: (2026)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
por: Tharwani, Jay, et al.
Publicado: (2025)
por: Tharwani, Jay, et al.
Publicado: (2025)
Performance optimization of BLAS algorithms with band matrices for RISC-V processors
por: Pirova, Anna, et al.
Publicado: (2025)
por: Pirova, Anna, et al.
Publicado: (2025)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
por: Accordi, Gianmarco, et al.
Publicado: (2025)
por: Accordi, Gianmarco, et al.
Publicado: (2025)
RAID Organizations for Improved Reliability and Performance: A Not Entirely Unbiased Tutorial
por: Thomasian, Alexander
Publicado: (2023)
por: Thomasian, Alexander
Publicado: (2023)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026)
por: Jain, Rutwik, et al.
Publicado: (2026)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
por: Landau, Diogo, et al.
Publicado: (2026)
por: Landau, Diogo, et al.
Publicado: (2026)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
por: Tharwani, Jay, et al.
Publicado: (2024)
por: Tharwani, Jay, et al.
Publicado: (2024)
Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM
por: Ren, Xuanzhengbo, et al.
Publicado: (2026)
por: Ren, Xuanzhengbo, et al.
Publicado: (2026)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
por: Takahashi, Keichi, et al.
Publicado: (2023)
por: Takahashi, Keichi, et al.
Publicado: (2023)
Experimental Assessment of Containers Running on Top of Virtual Machines
por: Aqasizade, Hossein, et al.
Publicado: (2024)
por: Aqasizade, Hossein, et al.
Publicado: (2024)
Ejemplares similares
-
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
por: Katagiri, Takahiro, et al.
Publicado: (2024) -
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
por: Apanasevich, L., et al.
Publicado: (2024) -
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
por: Bogart, Christopher, et al.
Publicado: (2025) -
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
por: Papavasileiou, Ioannis, et al.
Publicado: (2026) -
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
por: Debnath, Shimul, et al.
Publicado: (2026)