TALP-Pages: An easy-to-integrate continuous performance monitoring framework

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Seitz, Valentin, Trilaksono, Jordy, Garcia-Gasulla, Marta
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917012117127168
author Seitz, Valentin
Trilaksono, Jordy
Garcia-Gasulla, Marta
author_facet Seitz, Valentin
Trilaksono, Jordy
Garcia-Gasulla, Marta
contents Ensuring good performance is a key aspect in the development of codes that target HPC machines. As these codes are under active development, the necessity to detect performance degradation early in the development process becomes apparent. In addition, having meaningful insight into application scaling behavior tightly coupled to the development workflow is helpful. In this paper, we introduce TALP-Pages, an easy-to-integrate framework that enables developers to get fast and in-repository feedback about their code performance using established fundamental performance and scaling factors. The framework relies on TALP, which enables the on-the-fly collection of these metrics. Based on a folder structure suited for CI which contains the files generated by TALP, TALP-Pages generates an HTML report with visualizations of the performance factor regression as well as scaling-efficiency tables. We compare TALP-Pages to tracing-based tools in terms of overhead and post-processing requirements and find that TALP-Pages can produce the scaling-efficiency tables faster and under tighter resource constraints. To showcase the ease of use and effectiveness of this approach, we extend the current CI setup of GENE-X with only minimal changes required and showcase the ability to detect and explain a performance improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12436
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TALP-Pages: An easy-to-integrate continuous performance monitoring framework
Seitz, Valentin
Trilaksono, Jordy
Garcia-Gasulla, Marta
Distributed, Parallel, and Cluster Computing
Performance
Ensuring good performance is a key aspect in the development of codes that target HPC machines. As these codes are under active development, the necessity to detect performance degradation early in the development process becomes apparent. In addition, having meaningful insight into application scaling behavior tightly coupled to the development workflow is helpful. In this paper, we introduce TALP-Pages, an easy-to-integrate framework that enables developers to get fast and in-repository feedback about their code performance using established fundamental performance and scaling factors. The framework relies on TALP, which enables the on-the-fly collection of these metrics. Based on a folder structure suited for CI which contains the files generated by TALP, TALP-Pages generates an HTML report with visualizations of the performance factor regression as well as scaling-efficiency tables. We compare TALP-Pages to tracing-based tools in terms of overhead and post-processing requirements and find that TALP-Pages can produce the scaling-efficiency tables faster and under tighter resource constraints. To showcase the ease of use and effectiveness of this approach, we extend the current CI setup of GENE-X with only minimal changes required and showcase the ability to detect and explain a performance improvement.
title TALP-Pages: An easy-to-integrate continuous performance monitoring framework
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2510.12436