Where did we fail? -- Reproducing build failures in embedded open source software

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fu, Han, Ermedahl, Andreas, Eldh, Sigrid, Wiklund, Kristian, Haller, Philipp, Artho, Cyrille
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917448400240640
author Fu, Han
Ermedahl, Andreas
Eldh, Sigrid
Wiklund, Kristian
Haller, Philipp
Artho, Cyrille
author_facet Fu, Han
Ermedahl, Andreas
Eldh, Sigrid
Wiklund, Kristian
Haller, Philipp
Artho, Cyrille
contents Due to hardware-software co-development in embedded systems, continuous integration (CI) builds frequently fail because of complex cross-compilation, board configurations, and toolchain constraints. Although CI build logs contain valuable diagnostic information, they are short-lived and difficult to reuse due to heterogeneous runners, toolchains, and log formats. To address these challenges, we present PhantomRun, a unified abstraction layer and publicly reusable dataset that standardizes the retrieval, storage, and reproduction of CI build logs and metadata. Across 4628 failing CI runs, we reconstructed 91.8% of builds and preserved execution outcomes in 98% of evaluated cases. PhantomRun provides two core capabilities: retrieving the build log of any commit and faithfully re-executing the corresponding build in a controlled environment. By exposing all build artifacts and metadata in a uniform, machine-readable format, PhantomRun enables reproducible and longitudinal studies of CI failures. An empirical evaluation shows that reproduced builds closely match their originals, typically differing only in timestamps or minor nondeterministic reordering, demonstrating the feasibility of large-scale historical CI reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27075
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Where did we fail? -- Reproducing build failures in embedded open source software
Fu, Han
Ermedahl, Andreas
Eldh, Sigrid
Wiklund, Kristian
Haller, Philipp
Artho, Cyrille
Software Engineering
Due to hardware-software co-development in embedded systems, continuous integration (CI) builds frequently fail because of complex cross-compilation, board configurations, and toolchain constraints. Although CI build logs contain valuable diagnostic information, they are short-lived and difficult to reuse due to heterogeneous runners, toolchains, and log formats. To address these challenges, we present PhantomRun, a unified abstraction layer and publicly reusable dataset that standardizes the retrieval, storage, and reproduction of CI build logs and metadata. Across 4628 failing CI runs, we reconstructed 91.8% of builds and preserved execution outcomes in 98% of evaluated cases. PhantomRun provides two core capabilities: retrieving the build log of any commit and faithfully re-executing the corresponding build in a controlled environment. By exposing all build artifacts and metadata in a uniform, machine-readable format, PhantomRun enables reproducible and longitudinal studies of CI failures. An empirical evaluation shows that reproduced builds closely match their originals, typically differing only in timestamps or minor nondeterministic reordering, demonstrating the feasibility of large-scale historical CI reconstruction.
title Where did we fail? -- Reproducing build failures in embedded open source software
topic Software Engineering
url https://arxiv.org/abs/2604.27075