Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Mandel, Torsten, Bader, Jonathan, Yoo, Hanyoung, Kraft, Stephan
Format:	Preprint
Published:	2026
Subjects:	Distributed, Parallel, and Cluster Computing
Online Access:	https://arxiv.org/abs/2604.12673
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866910128844832768
author	Mandel, Torsten Bader, Jonathan Yoo, Hanyoung Kraft, Stephan
author_facet	Mandel, Torsten Bader, Jonathan Yoo, Hanyoung Kraft, Stephan
contents	Large enterprises often operate extensive Continuous Integration (CI) pipelines on large, heterogeneous compute clusters, where conservative, statically defined resource requirements are used to ensure build reliability. This practice leads to substantial system memory over-allocation, reduced cluster utilization, and increased operational costs. In this paper, we motivate the need for intelligent resource prediction by analyzing over 300,000 historical build executions from a production CI environment with more than one thousand compute nodes. Our analysis shows that, on average, more than 60% of allocated system memory remains unused. We then compare multiple machine learning approaches for predicting build task memory usage, including classification-based methods and regression-based quantile prediction. Our final solution employs a LightGBM-XGBoost quantile regression ensemble optimized to minimize under-allocation while reducing over-provisioning. We integrate this solution into the production CI pipeline via a microservice-based orchestration layer, achieving average memory savings of approximately 36GB per build and reducing under-allocation rates to below 0.3% without negatively impacting build execution times.
format	Preprint
id	arxiv_https___arxiv_org_abs_2604_12673
institution	arXiv
publishDate	2026
record_format	arxiv
spellingShingle	Intelligent resource prediction for SAP HANA continuous integration build workloads Mandel, Torsten Bader, Jonathan Yoo, Hanyoung Kraft, Stephan Distributed, Parallel, and Cluster Computing Large enterprises often operate extensive Continuous Integration (CI) pipelines on large, heterogeneous compute clusters, where conservative, statically defined resource requirements are used to ensure build reliability. This practice leads to substantial system memory over-allocation, reduced cluster utilization, and increased operational costs. In this paper, we motivate the need for intelligent resource prediction by analyzing over 300,000 historical build executions from a production CI environment with more than one thousand compute nodes. Our analysis shows that, on average, more than 60% of allocated system memory remains unused. We then compare multiple machine learning approaches for predicting build task memory usage, including classification-based methods and regression-based quantile prediction. Our final solution employs a LightGBM-XGBoost quantile regression ensemble optimized to minimize under-allocation while reducing over-provisioning. We integrate this solution into the production CI pipeline via a microservice-based orchestration layer, achieving average memory savings of approximately 36GB per build and reducing under-allocation rates to below 0.3% without negatively impacting build execution times.
title	Intelligent resource prediction for SAP HANA continuous integration build workloads
topic	Distributed, Parallel, and Cluster Computing
url	https://arxiv.org/abs/2604.12673

Similar Items