Learning From Software Failures: A Case Study at a National Space Research Center

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anandayuvaraj, Dharun, Singla, Tanmay, Hammadeh, Zain, Lund, Andreas, Holloway, Alexandra, Davis, James C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917294615035904
author Anandayuvaraj, Dharun
Singla, Tanmay
Hammadeh, Zain
Lund, Andreas
Holloway, Alexandra
Davis, James C.
author_facet Anandayuvaraj, Dharun
Singla, Tanmay
Hammadeh, Zain
Lund, Andreas
Holloway, Alexandra
Davis, James C.
contents Software failures can have significant consequences, making learning from failures a critical aspect of software engineering. While software organizations are recommended to conduct postmortems, the effectiveness and adoption of these practices vary widely. Understanding how engineers gather, document, share, and apply lessons from failures is essential for improving reliability and preventing recurrence. High-reliability organizations (HROs) often develop software systems where failures carry catastrophic risks, requiring continuous learning to ensure reliability. These organizations provide a valuable setting to examine practices and challenges for learning from software failures. Such insight could help develop processes and tools to improve reliability and prevent recurrence. However, we lack in-depth industry perspectives on the practices and challenges of learning from failures. To address this gap, we conducted a case study through 10 in-depth interviews with research software engineers at a national space research center. We examine how they learn from failures: how they gather, document, share, and apply lessons. To assess transferability, we include data from 5 additional interviews at other HROs. Our findings provide insight into how engineers learn from failures in practice. To summarize: (1) failure learning is informal, ad hoc, and inconsistently integrated into SDLC; (2) recurring failures persist due to absence of structured processes; and (3) key challenges, including time constraints, knowledge loss from turnover and fragmented documentation, and weak process enforcement, undermine systematic learning. Our findings deepen understanding of how software engineers learn from failures and offer guidance for improving failure management practices.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning From Software Failures: A Case Study at a National Space Research Center
Anandayuvaraj, Dharun
Singla, Tanmay
Hammadeh, Zain
Lund, Andreas
Holloway, Alexandra
Davis, James C.
Software Engineering
Software failures can have significant consequences, making learning from failures a critical aspect of software engineering. While software organizations are recommended to conduct postmortems, the effectiveness and adoption of these practices vary widely. Understanding how engineers gather, document, share, and apply lessons from failures is essential for improving reliability and preventing recurrence. High-reliability organizations (HROs) often develop software systems where failures carry catastrophic risks, requiring continuous learning to ensure reliability. These organizations provide a valuable setting to examine practices and challenges for learning from software failures. Such insight could help develop processes and tools to improve reliability and prevent recurrence. However, we lack in-depth industry perspectives on the practices and challenges of learning from failures. To address this gap, we conducted a case study through 10 in-depth interviews with research software engineers at a national space research center. We examine how they learn from failures: how they gather, document, share, and apply lessons. To assess transferability, we include data from 5 additional interviews at other HROs. Our findings provide insight into how engineers learn from failures in practice. To summarize: (1) failure learning is informal, ad hoc, and inconsistently integrated into SDLC; (2) recurring failures persist due to absence of structured processes; and (3) key challenges, including time constraints, knowledge loss from turnover and fragmented documentation, and weak process enforcement, undermine systematic learning. Our findings deepen understanding of how software engineers learn from failures and offer guidance for improving failure management practices.
title Learning From Software Failures: A Case Study at a National Space Research Center
topic Software Engineering
url https://arxiv.org/abs/2509.06301