Optimisation of ATLAS computing resource usage through a modern HEP Benchmark Suite via HammerCloud and Big PanDA

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Szczepanek, Natalia, Giordano, Domenico, Glushkov, Ivan, Borge, Gonzalo Menendez, Di Girolamo, Alessandro, Lory, Alexander, Vukotic, Ilija
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916603046658048
author Szczepanek, Natalia
Giordano, Domenico
Glushkov, Ivan
Borge, Gonzalo Menendez
Di Girolamo, Alessandro
Lory, Alexander
Vukotic, Ilija
author_facet Szczepanek, Natalia
Giordano, Domenico
Glushkov, Ivan
Borge, Gonzalo Menendez
Di Girolamo, Alessandro
Lory, Alexander
Vukotic, Ilija
contents In April 2023, HEPScore23, the new benchmark based on HEP specific applications, was adopted by WLCG, replacing HEP-SPEC06. As part of the transition to the new benchmark, the CPU corepower published by the sites needed to be compared with the effective power observed while running ATLAS workloads. One aim was to verify the conversion rate between the scores of the old and the new benchmark. The other objective was to understand how the HEPScore performs when run on multi-core job slots, so exactly like the computing sites are being used in the production environment. Our study leverages the HammerCloud infrastructure and the PanDA Workload Management System to collect a large benchmark statistic across 136 computing sites using an enhanced HEP Benchmark Suite. It allows us to collect not only performance metrics, but, thanks to plugins, it also collects information such as machine load, memory usage and other user-defined metrics during the execution and stores it in an OpenSearch database. These extensive tests allow for an in-depth analysis of the actual, versus declared computing capabilities of these sites. The results provide valuable insights into the real-world performance of computing resources pledged to ATLAS, identifying areas for improvement while spotlighting sites that underperform or exceed expectations. Moreover, this helps to ensure efficient operational practices across sites. The collected metrics allowed us to detect and fix configuration issues and therefore improve the experienced performance.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04853
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimisation of ATLAS computing resource usage through a modern HEP Benchmark Suite via HammerCloud and Big PanDA
Szczepanek, Natalia
Giordano, Domenico
Glushkov, Ivan
Borge, Gonzalo Menendez
Di Girolamo, Alessandro
Lory, Alexander
Vukotic, Ilija
Distributed, Parallel, and Cluster Computing
High Energy Physics - Experiment
In April 2023, HEPScore23, the new benchmark based on HEP specific applications, was adopted by WLCG, replacing HEP-SPEC06. As part of the transition to the new benchmark, the CPU corepower published by the sites needed to be compared with the effective power observed while running ATLAS workloads. One aim was to verify the conversion rate between the scores of the old and the new benchmark. The other objective was to understand how the HEPScore performs when run on multi-core job slots, so exactly like the computing sites are being used in the production environment. Our study leverages the HammerCloud infrastructure and the PanDA Workload Management System to collect a large benchmark statistic across 136 computing sites using an enhanced HEP Benchmark Suite. It allows us to collect not only performance metrics, but, thanks to plugins, it also collects information such as machine load, memory usage and other user-defined metrics during the execution and stores it in an OpenSearch database. These extensive tests allow for an in-depth analysis of the actual, versus declared computing capabilities of these sites. The results provide valuable insights into the real-world performance of computing resources pledged to ATLAS, identifying areas for improvement while spotlighting sites that underperform or exceed expectations. Moreover, this helps to ensure efficient operational practices across sites. The collected metrics allowed us to detect and fix configuration issues and therefore improve the experienced performance.
title Optimisation of ATLAS computing resource usage through a modern HEP Benchmark Suite via HammerCloud and Big PanDA
topic Distributed, Parallel, and Cluster Computing
High Energy Physics - Experiment
url https://arxiv.org/abs/2502.04853