scieee AI-readable full text Open interactive document viewer

Artifact: Learning from software failures: A case study at a national space agency

Anandayuvaraj, Dharun

Abstract

Artifact for "Learning from software failures: A case study at a national space agency"

Full text

ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil Dharun Anandayuvaraj, Tanmay Singla, Zain A. H. Hammadeh, Andreas Lund, Alexandra Holloway, and James C. Davis Outline of Appendices The appendix contains the following material: •Appendix A: The interview protocol. • Appendix B: The full code book organized by themes & subjects. • Appendix C: The evolution of the Codebook used in our analysis. • Appendix D: Additional demographics for the case study subjects. A Interview Protocol Table 1 gave a summary of the interview protocol. Here we describe the full protocol in Table 4. Table 5 maps our interview protocol to our research questions. B Code Categorized by Themes Table 6 presents the full set of codes identified through our qualitative analysis. We organized these codes into three overarching themes that emerged from the data: (1) failure learning is informal and inconsistent, (2) recurring failures and the illusion of effective failure management, and (3) challenges within the learning-fromfailures process. Within each theme, we group related codes into subthemes to highlight more specific patterns. For each code, we also indicate which subjects (S1–S10, PS1–PS4, PTeam) mentioned the issue during interviews. This structure allows us to both quantify the spread of each code and contextualize how failure-related practices differ across participants. C Code Evolution Diagram Figure 4 describes the evolution of our codebook over six rounds of revisions. D Additional Demographics for Case Study Subjects Figure 5 illustrates the highest degree held by participants. Figure 6 illustrates the disciplinary background of participants. Figure 7 illustrates where participants reported learning about software engineering practices. Learning From Software Failures: A Case Study at a National Space Research Center ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil Table 4: Interview Protocol for Postmortems in Software Failures Category Questions Reflective Example Q1: Can you provide an example of a time when your team studied a software/engineering issue to gather lessons? Q2: After the first time, did this issue reoccur? Q3: Were measures taken to prevent the issue from reoccurring? How? Q4: Did the issue reoccur after these measures? If so, what actions were taken afterward? Q5: If recurred: Can you provide an example of an issue with long-term remediation where the issue stopped recurring? Q6: If did not recur: Can you provide an example of an issue recurring despite initial fixes? Postmortem Analysis Q7: How does your team study or handle software issues when they occur? Q8: What happens once the issue is fixed? Q9: In what scenarios are issues studied beyond just fixing the issue? Q10: Are these issues communicated amongst other team members or stakeholders? How? Q11: Would your team prioritize studying issues with impacts on humans or the environment? Q12: Do you have insights on how reflective practices differ for hardware or mechanical failures? Q13: Are there meetings to discuss lessons from issues? How are they structured? Q14: How are issues and lessons documented? How often and in what cases are they referenced? Q15: How are these issues and lessons stored (Internal wiki, Jira, GitLab, Confluence)? Learning Integration Q16: How are lessons learned integrated into future product development? Q17: Are issues from past projects used in fault modeling during new projects? Q18: If not used, would it be beneficial to use past lessons? Q19: If so, how are they beneficial and effective? Q20: What changes resulted from studying past issues? Q21: How are issues and lessons shared across disciplines or teams? Q22: How are these lessons communicated to new team members (training, reference material)? Processes Q23: Regarding activities you’ve described, how standardized are these processes? Q24: Are these standardized across all teams? Why or why not? Q25: Were these activities different in your past teams/organizations? How? Q26: Do you see gaps or opportunities for improvement in these processes? Q27: To what extent are activities automated, from documenting to learning from issues? Q28: What further tools or automation would you suggest to help document and learn from issues? Benefits of Postmortems Q29: Textbooks and research literature say that studying and learning from failures is an important process. Do you believe that your team would/has benefit from such processes? Why? Q30: Do you think there would be changes in your products due to such processes? Q31: Would formal processes like reading clubs or role-playing exercises be useful? Why? Q32: Do you think these processes are practical? What constraints concern you? Automation Q33: What type of automation would help document issues beyond just tracking them? Q34: What type of automation would help learn from issues? Q35: Would a database linking issues with lessons learned help your development process? Q36: How would a chatbot for querying project-specific issue information assist your team? Table 5: Interview Protocol organized by Research Questions (RQs). Some interview questions address more than one research question. RQ Questions RQ1: How do software teams gather lessons from failures? Q1–Q6 (Reflective examples of recurring or remediated issues) Q7–Q9 (How issues are studied and handled when they occur) Q13–Q15 (Meetings to discuss lessons; how lessons are documented and stored) RQ2: How are these lessons shared & integrated into the SDLC? Q10 (Communication with team members or stakeholders) Q16–Q22 (Integration into future projects, fault modeling, benefits, onboarding/- training, cross-team sharing) RQ3: What are the challenges in learning from failures? Q11–Q12 (Prioritization of issues; differences from hardware/mechanical failures) Q23–Q25 (Standardization across teams and organizations; comparison with past teams/orgs) Q26–Q32 (Gaps, opportunities for improvement, practicality, constraints, perceived benefits) Q33–Q36 (Automation and tooling challenges/opportunities) ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil Dharun Anandayuvaraj, Tanmay Singla, Zain A. H. Hammadeh, Andreas Lund, Alexandra Holloway, and James C. Davis Table 6: The codebook developed from and applied to our data, organized by themes and our subjects (outlined in Table 2). Code Subjects Theme 1: Failure learning is informal & inconsistent Gathering lessons from failures is informal & inconsistent Example of ad hoc discussion of failure S2, S5, S6, S7, S10 Failure discussion meeting structure is ad hoc S1, S2, S3, S4, S5, S6, S8, S9, S10, PS2, PS3, PTeam There is a not a blaming culture, enabling open failure discussion S3, S5, S6 There is no process to learn from failures S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, PS1, PS3, PTeam Studying, learning, sharing failure knowledge is ad hoc S1, S2, S4, S5, S6, S7, S8, S9, PS2, PTeam Reflection of good, bad, and improvements ad hoc S3, S4, S7, S8, S10 Lessons learnt are gather for project management but not software S7, S9 Lessons learnt is gathered ad hoc S2, S5, S6, S7, S8, S9, PS2 Example of ad hoc gathering of lessons learnt S2, S3, S5, S6, S7 Disconnect between subjects thinking that: discussion of failure = gathering lessons learnt S2, S4, S6, S9, S10, PTeam Disconnect: assumption that lessons learnt are gathered when they are not S3, S4, S6, S7, S9, PS1, PS2, PS3, PTeam Failure is discussed upto initial fix ( prevention or lessons learnt) S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, PS2, PS3 Documentation of lessons learned is inconsistent & fragmented Example of ad hoc documentation of failure S2, S6 Failure knowledge documentation practices are ad hoc S1, S2, S3, S4, S5, S6, S7, S10, PS1, PS2, PS3 Methods to document failure knowledge S2, S4, S5, S7, S8, PS1, PS2, PS3, PTeam Failure is documented upto fix ( prevention or lessons learnt) S1, S2, S3, S4, S6, S7, S8, S9, S10, PS1, PS3, PTeam Lessons learnt is documented ad hoc S2, S6, S7, S8, S9, PTeam Lessons learnt is NOT explicitly documented S1, S2, S4, S5, S6, S8, S10, PS1 Lessons learned are transferred & applied informally Failure knowledge is shared ad hoc S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, PTeam Past failure knowledge is provided to new members S2, S4, PS1, PS2 Past failure knowledge is NOT provided to new members S1, S5, S9, PS3, PTeam Senior members transfer internalized lessons ad hoc S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, PS1 Sharing is complicated with classified knowledge S2, S5, S6, S7, S8, S10 Lessons learnt is NOT standardly shared S1, S2, S10, PTeam Example of lessons learnt being applied in practice S1, S3, S6, S7, S8, S9, S10, PS2, PTeam Disconnect between past failure knowledge being referenced when it is not referenced often S2, S3, S4, S7, S9, PS1 Failure knowledge documentation is reviewed later S2, S3, S4, S5, S8, S10, PS2, PTeam Failure knowledge documentation is NOT reviewed later S1, S2, S3, S6, S7, S8, S10, PS2, PS3 Theme 2: Recurring failures and the illusion of effective failure management Disconnect between subjects thinking their failure management is working, except they report recurring issues S2, S7, PS2 Disconnect: assumption that general software practices (reusability, code quality, etc) can prevent recurring failures S1, S3, S4, S6, S9, PS3 There are recurring failures S1, S4, S5, S6, S7, S9, PS3, PTeam Example of failure not recurring S2, S3, S6, S7 Example of failure recurring S1, S2, S3, S5, S6, S7, S8, S10, PS3 Recurring failure due to inadequate failure documentation practices S1, S2, S7 Recurring failure due to not sharing S2, S7 Recurring failure due to fix for a specific case (Need generalizable processes, postmortems) S2 Disconnect: assumption that only new members need process to learn from failures S7, PS2 Theme 3: Challenges within the learning-from-failures process Formal process for lessons learnt is standardized, but is too general and not useful S1, S7 SDLC is different/informal due to research org S1, S3, S4, S6, S8 Loss of knowledge due to dynamic team changes S2, S4, S5, S8, S9, PTeam Loss of knowledge due to inadequate failure documentation practices S2, S5, S7, S9, PS1 Lack of time and resources to conduct traditional postmortems S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, PS1, PS3, PTeam Inclination to have a process to learn from failures if impact was on physical world S3, S5, PS1, PS3 More effort to prevent failures in traditional engineering at earlier dev phase S1, S4, S5, S6, S7, S9, PTeam Learning From Software Failures: A Case Study at a National Space Research Center ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil Figure 4: Codebook Evolution: The codebook underwent six rounds of revisions, with changes highlighted by different colors: round 2 (blue), round 3 (red), and round 6 (purple). Arrows indicate the replacement of codes. The final codes from round 6 were used in this paper, and categorized thematically. ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil Dharun Anandayuvaraj, Tanmay Singla, Zain A. H. Hammadeh, Andreas Lund, Alexandra Holloway, and James C. Davis Figure 5: Highest degree held by participants. Figure 6: Disciplinary background of participants. Figure 7: Where participants reported learning about software engineering practices.