D4.7 - Operational plans for the pan-European umbrella ecosystem
Abstract
This deliverable defines the context and rationale for the existence of the pan-European ecosystem, its governance model (structure, guiding principles, incentives of partners), specific operations, and rules of its operational activities as a pan-European federated learning ecosystem (for ongoing and new members, for the submission of use cases and use of federated analysis infrastructures). These aspects define the path to generating value, insights, impact, and sustainability through collaboration and federated learning, with iterative tailoring to the stakeholders within the STRONG-AYA Consortium.
Full text
1 A new, interdisciplinary, multi-stakeholder European network to improve healthcare services, research and outcomes for Adolescents and Young Adults with cancer.
STRONG-AYA – No. 101057482 – D4.7 2 Deliverable Report WP4 – Operation of STRONG-AYA ecosystem, stakeholder and patient involvement, dissemination, exploitation, communication Deliverable D4.7 Operational plans for the pan-European umbrella ecosystem Due date of deliverable: 30/09/2023 Actual submission date: 22/03/2024 Revised from Consolidated Report of 25-07-2024, resubmitted 30-09-2024 Project: STRONG-AYA Lead Contributor Oana Lindner (UOL) Email [email protected] Other Contributors Dan Stark, Oana Lindner, Emily Connearn, Richard Feltbower, Nicola Hughes (UOL); Nicole Collaco (UOS); Leonard Wee, Marine Jacquemin, Joshi Hogenboom, Flora Lysen, Darian Meacham (UNIMAAS); Simone Hanebaum (NKI); Urska Kosir (PAB). Emails [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; marine[email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]. Due date 30/09/2023 Delivery date 22/03/2024 Deliverable type R Dissemination level PU
STRONG-AYA – No. 101057482 – D4.7 3 Description of Work Version Date First draft for review V1.0 22/03/2024 Redraft from review V2.0 27/09/2024 Description: Operational plan containing description of main actions required for the deployment and functioning of the pan-European ecosystem 5. Publishable summary (max ½ page) This deliverable defines the context and rationale for the existence of the pan-European ecosystem, its governance model (structure, guiding principles, incentives of partners), specific operations, and rules of its operational activities as a pan-European federated learning ecosystem (for ongoing and new members, for the submission of use cases and use of federated analysis infrastructures). These aspects define the path to generating value, insights, impact, and sustainability through collaboration and federated learning, with iterative tailoring to the stakeholders within the STRONG-AYA Consortium. This deliverable is part of WP4, defining the actions that all institutions need to fulfil for the operation of a sustainable pan-European STRONG-AYA ecosystem. It builds upon the existing national/local ecosystems within STRONG-AYA and previous documents produced within the Consortium related to the architecture of the ecosystem (D2.1), code of conduct of members (D2.2), data management (D3.1), technical blueprint (D3.3), previously-defined local operational plans (D4.4, D4.6) and definition of use cases (D4.8), and key performance indicators (D5.2). Apart from these documents, which all partners are aware of, the current report should be read in conjunction with the ROPA/DPIA and Grant agreement for any additional details. Details encompassed by the actions defined within the overarching national ecosystem operational plan (D4.4) and operational plans for national ecosystems (D4.6) feed directly into this document, which brings together the actions and operations for the development of national infrastructures feeding into our panEuropean umbrella ecosystem. D4.4 and D.4.6 both rely on interviews and discussions with each national partner, evaluating the state of preparedness to build from local to national ecosystems. The pan-European ecosystem described here is reliant on each local centre working nationally, to create capacity to build towards a population-based national StTRONG-AYA ecosystem, which can communicate with other nations. Consistent with the other deliverables to which this report is linked, it aims to be a living document, developed iteratively throughout the project and beyond. It will evolve as new partners and interested stakeholders may join the Consortium, may contribute data, and the landscape of governance, ethics, and security changes in time at international levels.
STRONG-AYA – No. 101057482 – D4.7 4 6. Table of Contents PUBLISHABLE SUMMARY (MAX ½ PAGE) .................................................................................................................... 3 TABLE OF CONTENTS .................................................................................................................................................. 4 DEFINITIONS ............................................................................................................................................................... 6 ABBREVIATIONS ......................................................................................................................................................... 7 4. EXECUTIVE SUMMARY ............................................................................................................................................ 8 5. INTRODUCTION ...................................................................................................................................................... 8 5.1. PROJECT BACKGROUND ........................................................................................................................................... 8 5.2. DEFINITION OF THE PAN-EUROPEAN ECOSYSTEM ........................................................................................................ 10 5.3. IMPLEMENTATION OF FEDERATED LEARNING ............................................................................................................. 11 6. GOVERNANCE MODEL .......................................................................................................................................13 6.1. STRUCTURE OF THE CONSORTIUM ........................................................................................................................... 14 6.1.1. Existing Consortium partners.................................................................................................................. 14 6.1.2. Committee for Coordination ................................................................................................................... 14 6.2. FOUNDING PRINCIPLES .......................................................................................................................................... 14 6.2.1. Value-based healthcare .......................................................................................................................... 14 6.2.2. FAIR principles ........................................................................................................................................ 15 6.2.3. Open science principles ........................................................................................................................... 15 6.2.4. Transparency .......................................................................................................................................... 15 6.2.5. Standardisation ...................................................................................................................................... 17 6.2.6. Inclusive and focused on patient benefit ................................................................................................ 17 6.2.7. Privacy and confidentiality-preserving ................................................................................................... 18 6.2.8. Scalable and sustainable ........................................................................................................................ 19 6.3. ALIGNING INCENTIVES AND MOTIVES ACROSS CONSORTIUM MEMBERS ........................................................................... 20 6.4. MANAGING NATIONAL AND INSTITUTIONAL DIFFERENCES ............................................................................................. 21 7. STANDARDS AND RULES FOR THE FEDERATED LEARNING ECOSYSTEM .............................................................26 7.1. RULES FOR ONGOING AND NEW PARTNERS ................................................................................................................ 27 7.2. APPLICATION REQUIREMENTS ................................................................................................................................. 27 7.3. DECISION PROCESS ............................................................................................................................................... 28 7.4. MANAGEMENT OF USER ACCESS AND USAGE .............................................................................................................. 28 7.5. RESOLUTION OF DISPUTES AND AMBIGUOUS SITUATIONS .............................................................................................. 32 7.6. SUBMISSION OF USE CASES ..................................................................................................................................... 33 8. RULES ON USING THE ECOSYSTEM FOR FEDERATED LEARNING AND ANALYTICS ..............................................34 8.1. ACTIVE CONTRIBUTION TO THE DATA ECOSYSTEM ........................................................................................................ 34 8.2. ABIDING BY THE TECHNICAL STANDARDS ................................................................................................................... 35 8.3. PRIVACY PRESERVATION ......................................................................................................................................... 35 8.4. SUPPORTING INTERPRETABILITY AND INTEROPERABILITY ............................................................................................... 35 8.5. TRACKING SUCCESS WITH KPIS ................................................................................................................................ 36 8.6. USING THE INFRASTRUCTURE .................................................................................................................................. 37 8.7. INTELLECTUAL PROPERTY GUIDELINES ....................................................................................................................... 41 8.8. SHARING OF BENEFITS - COMMUNICATION AND DISSEMINATION OF OUTPUTS .................................................................. 41
STRONG-AYA – No. 101057482 – D4.7 5 9. RE-AUDITING OF RULES AND STANDARDS ......................................................................................... 44 10. COMMITMENTS OF CONSORTIUM AND PARTNERS ......................................................................................44 11. CONCLUSION .................................................................................................................................................45
STRONG-AYA – No. 101057482 – D4.7 6 7. Definitions STRONG AYA consortium members are referred to as following within this text: 1. NKI-AVL – Stichting het Nederlands Kanker Instituut – Antoni van Leeuwenhoek Ziekenhuis (NL) 2. YCE – Youth Cancer Europe (RO) 3. INT – Fondazione IRCCS Instituto Nazionale dei Tumori (IT) 4. FFUND – FFUND BV (NL) 5. CLB – Centre de Lutte Contre le Cancer Leon Berard (FR) 6. ECO – European Cancer Organisation (BE) 7. UNIMAAS – Universiteit Maastricht (NL) 8. IKNL – Stichting Integraal Kankercentrum Nederland (NL) 9. EORTC – European Organisation for Research and Treatment of Cancer AISBL (BE) 10. IGR – Institut Gustave Roussy (FR) 11. MSCNRIO – Narodowy Instytut Onkologii im. Marii Sklodowskiej-Curie – Panstwowy Instytut Badawczy (Marie Sklodowska-Curie National Research Institute of Oncology) (PL) 12. UOM – University of Manchester (UK) 13. UOL – University of Leeds (UK) 14. LTHT – Leeds Teaching Hospitals National Health Service Trust (UL) 15. SOUTHAMPTON – University of Southampton (UK) Grant Agreement (including its annexes and amendments): the agreement signed between the beneficiaries of the HORIZON Research and Innovations Actions (hereafter referred to as Horizon) and the European Health and Digital Executive Agency (hereafter referred to as HADEA) for the undertaking of the STRONG AYA project (Grant Agreement no. 101057482). Beneficiary: Signatories of the Grant Agreement Associated Partner: Entities which participate in the action but without the right to charge costs or claim contributions. Project: the sum of all activities carried out in the framework of the Grant Agreement. Consortium: the STRONG AYA consortium, including all the aforementioned partners. Consortium Agreement: The agreement made between STRONG AYA members for the implementation and execution of the action outlined in the Grant Agreement. The agreement shall not affect the parties’ obligations to HADEA on behalf of the European Union, and/or to one another arising from the Grant Agreement.
STRONG-AYA – No. 101057482 – D4.7 7 8. Abbreviations Acronym/Abbreviation Meaning API Application Programming Interface AYA Adolescent and Young Adult HCP Health Care Provider PRO Patient Reported Outcome PROM Patient Reported Outcome Measure CfC Committee for Coordination COS Core Outcome Set FAIR Findable, Accessible, Interoperable, Reusable FL Federated learning GDPR General Data Protection Regulation MDW Medical Data Works PHT Personal Health Train PI Principal Investigator PLUTO Public Value Assessment Tool WP Work Package WPL Work Package Lead(s) WP1 Work Package 1 (Development Core Outcome Set AYA with cancer & data collection) WP2 Work Package 2 (Governance, Data Security and Ethics) WP3 Work Package 3 (Infrastructure and Interoperability) WP4 Work Package 4 (Operation of STRONG AYA ecosystems, stakeholder and patient involvement, dissemination, exploitation, communication) WP5 Work Package 5 (Scientific coordination and project management) KPI Key Performance Indicator OA Open Access PAB Patient Advisory Board EC European Commission HADEA European Health and Digital Executive Agency
STRONG-AYA – No. 101057482 – D4.7 8 9. 4. Executive summary 1. AYA cancers are rare but can be disabling for a large proportion of surviving young people; moreover, clinical and psychosocial research and care for AYA is not evenly distributed and developed across Europe. 2. Because they are rare, conducting adequately powered research that improves clinical and psychosocial care, and implementing research findings, while meeting young people where they are (i.e. in case they move geographically), requires a collaborative effort across Europe. 3. Data sharing to make informed decisions and conclusions about the care of AYA is not possible due to a number of challenges – including risks to data privacy and protection and different ways of applying ethical principles to data collection, management and sharing over time and between countries. 4. The STRONG-AYA Consortium aims to build a sustainable pan-European ecosystem that addresses some of these challenges through a defined set of principles, rules for joining the ecosystem and making data available for privacy-preserving federated learning, for the benefit of AYA care across Europe. 5. This document outlines the governing principles of the pan-European ecosystem to be adhered to by current and new members, the process of becoming a member, and the rules for contributing to, benefiting from, and ensuring the sustainability of a growing federated learning infrastructure. It also lays out the limitations, facilitators, risks and mitigations for each element of our panEuropean ecosystem, and the next methods and metrics required for further progress 10. 5. Introduction 10.1. Project background Adolescents and young adults (AYAs) with cancer form a unique group; they face age-specific issues and decreased quality of life. Cancer in AYAs aged 15-39 is rare, although 4-6 times more frequent than paediatric cancer (i.e. prepubescent period). However, this rarity does not reflect the significant personal and societal costs of cancer in this population, as reflected in the potential years of life lost or saved, the decreased productivity and quality of life and the life-long complications or disabilities1. Improving survival in AYA is more challenging than for children and older cancer survivors. This is partly due to the excess risk of second primary malignant neoplasms identified in this group compared to cohorts of younger and older patients2. Moreover, AYAs face some distinct challenges compared to older and younger counterparts in the treatment of their primary or secondary malignancies - a unique spectrum of cancer types, different tumour biology, unique complex psychological needs, distinct late sequelae, including impaired fertility, and palliative care3,4. 1 Stoneham SJ. AYA survivorship: The next challenge. Cancer 2020; 126: 2116-2119. 2 Keegan THM, Bleyer A, Rosenberg AS et al. Second Primary Malignant Neoplasms and Survival in Adolescent and Young Adult Cancer Survivors. JAMA Oncol 2017; 3: 1554-1557. 3 Bleyer A, Barr R, Hayes-Lattin B et al. The distinctive biology of cancer in adolescents and young adults. Nat Rev Cancer 2008; 8: 288-98. 4 Ferrari A, Stark D, Peccatori FA et al. Adolescents and young adults (AYA) with cancer: a position paper from the AYA Working Group of the European Society for Medical Oncology (ESMO) and the European Society for Paediatric Oncology (SIOPE). ESMO Open 2021; 6: 100096.
STRONG-AYA – No. 101057482 – D4.7 9 These consequences imply that the clinical management, treatment, diagnosis, and psychosocial support need to be designed and developed for the specific needs of AYA. For example, AYAs diagnosed with breast and prostate carcinomas have worse survival than older patients because of the biological differences, highlighting the need to target screening methodologies, treatments, and policies to their needs. AYAs with cancer also face significant psychological challenges, including substance abuse, mental health issues, suicidal ideations, and increased emotional burden from cancer and cancer-related morbidity. Finally, tailoring cancer care to AYAs’ needs is also made difficult by the low rate of participation by AYAs in clinical trials and cancer research5. Despite the increasing awareness and a growing body of the scientific literature, European health systems have not fully recognised and addressed the unique issues faced by AYAs with cancer, and therefore these individuals often do not receive the right type of support. In part, this may have resulted from the traditional dichotomy between the integrated paediatric (“patient/family-centred”) care services versus dispersed (“disease-centred”) adult oncology services4,6,7. As a result, up to half of AYAs with cancer report unmet informational and service needs, impacting both their direct (survival rates) and indirect (long-term effects and mental health) recovery and social integration8. Furthermore, aligned with this barrier is the low rate of health care utilisation by AYAs, especially primary care, in part due to low awareness of AYA malignancy among primary carers and, depending on the nation, due to challenges to sustain health insurance coverage5. According to a survey conducted through the networks of the AYA Working Group of the European Society for Medical Oncology (ESMO) and the European Society for Paediatric Oncology (SIOP Europe), 67% of practitioners do not have access to specialised centres for AYA with cancer, 67% had no access to a specialist cancer service for late effects management and 38% had no access to fertility specialists. Under-provision and inequality of AYA cancer care is common across Europe and especially in the Eastern and Southern-Eastern areas. Furthermore, the Working Group also reported an absence of outcome measures for monitoring and evaluating AYA cancer care programs9. Among the recommended future steps, one of the most important contributions to AYA research would be to pool data (e.g. patient-reported outcomes, clinical and treatment data) across institutions and countries and create large cohorts for researchers to address the burden of cancer in AYA10. However, there are challenges due to a lack of data standardisation, data interoperability, unitary (prospective) collection of outcomes of relevance for AYAs with cancer, data privacy and security concerns, as well as regional differences in the application of ethical principles governing data collection and sharing. The STRONG-AYA project aims to tackle the underrepresentation of AYA’s experiences and outcomes when navigating the healthcare system and in clinical care by developing national infrastructures for outcome data management and clinical decision-making within a pan-European ecosystem and establishing communication feedback for AYAs with cancer and their healthcare systems. This will be key to improving healthcare services, research, outcomes and policies for AYAs and to ultimately improved cancer care for this patient group. 5 Hayashi RJ. Adolescent and young adult cancer survivorship: The new frontier for investigation. Cancer 2019; 125: 1976-1978. 6 Osborn M, Johnson R, Thompson K et al. Models of care for adolescent and young adult cancer programs. Pediatr Blood Cancer 2019; 66: e27991. 7 Fardell JE, Patterson P, Wakefield CE et al. A Narrative Review of Models of Care for Adolescents and Young Adults with Cancer: Barriers and Recommendations. J Adolesc Young Adult Oncol 2018; 7: 148-152. 8 Keegan TH, Lichtensztajn DY, Kato I et al. Unmet adolescent and young adult cancer survivors information and service needs: a population-based cancer registry study. J Cancer Surviv 2012; 6: 239-250. 9 Saloustros E, Stark D, Michailidou K et al. Report on ESMO/SIOPE European Landscape project key results: Mapping the status and needs in AYA cancer care. Late-breaking and deferred publication abstracts public health 2017; 28, 5: V643. 10 Smith AW, Seibel NL, Lewis DR et al. Next steps for adolescent and young adult oncology workshop: An update on progress and recommendations for the future. Cancer 2016; 122: 988-999.
STRONG-AYA – No. 101057482 – D4.7 16 against a single Consortium member being the majority data owner which would create structural inequity and power conflicts at the decision making level. For the purpose of equity in data provision and data queries and hence to reduce misalignment resulting from power democracy, the STRONG-AYA pan-European ecosystem will ensure: Data stewardship and control will be jointly held by Principal investigators (PIs) and their teams of investigators. That is because focusing power only on the PI would remove the institutional investigators from the power structure and the data governance loop. Through joint power over data we ensure the context and subtlety related to data is maintained, and if data needs to be enriched, enhanced or requires additional management, this can be done in partnership within the local ecosystems, rather than it being the responsibility of a single person. The actions herein are not a new topology of ‘data control exercise’ by a single partner. That is, it is important that all partners are equal in terms of suggesting a use case (as long as it is congruent with the foundational data usage principles and incentives of the consortium) and all partners must have the same “power” in terms of being able to run an analysis and test their hypotheses. The consortium will ensure no one partner makes all the rules and no partner becomes a mere datafeeding station. “Joint controllership” according to GDPR – in terms of the data contributions by each partner and the analyses being run on each others datasets - becomes a central tenet of the legal basis for data usage and it is specified as part of the data protection agreement signed by consortium members. The federation must still be able to function (albeit in such cases on a reduced data sample) if one partner conscionably opts out of a particular analysis or use case. The actions and operations within the ecosystem are not ‘all-in or all-out’. For instance, it may be that a healthcare provider wants to be part of sarcoma studies but wants to opt out of breast cancer studies. From the STRONG-AYA consortium perspective, this complies with the principle of enshrined autonomy of each investigator in the consortium. This also respects local processes and obligations and regulations; for example sociodemographic data collection and analysis is permitted in the UK, but is illegal in France. The opt-out of French centres from such analyses is thus imperative and so maintaining a more flexible approach is an inclusive, rather than exclusive, structure. In the case of prospective (novel) data collection (such as PROs based on the COS), each local partner needs to ensure that transparency is present in communication with participants. This means that all participants have provided and documented informed consent, which contains the relevant information about data processing and data protection aligned to article 13 of the GDPR, including any local institutional guidelines and best practices. In the case of retrospective (historical) data available and provided by partners, article 4 of GDPR applies, thereby the data collected by a partner must meet all three grounds allowing the processing of such data. Namely, the data can be processed if a) the participant offered freely given, specific, and informed consent (Article 6 (1) (a)); b) processing is a legitimate interest exception such as there is a legitimate interest of controllers AND the legitimate interest cannot reasonably be achieved by using alternative means AND the interests and fundamental rights and freedom of the participant do not take precedence over the legitimate interests of the controller(s); c) the data is anonymised and therefore falls outside GDPR. As an example, a legitimate interest exception would be that of testing hypotheses for the purpose of public health. Finally, honoring this principle means that participants should find it easy to understand how specific information about their health is used or accessed. Participants should easily and reliably have control and be able to access their health data and they should be able to exercise these rights. Exactly what type of data
STRONG-AYA – No. 101057482 – D4.7 17 patients versus healthcare professionals will be able to see is decided by the CfC in collaboration with the entire consortium and the Patient Advisory Board to avoid undue distress or harmful decision making. 11.2.5. Standardisation of data One of the limitations identified in summarising and making valuable conclusions for clinical decisions in AYA’s clinical care internationally is the lack of data standardisation. This relates to the changing landscape of healthcare professionaland patient-reported outcomes that are considered relevant for the treatment, care, and research in AYA cancers. As AYAs are an under-researched demographic and straddle the traditional divisions between paediatric and adult oncology, the type of care an AYA experiences varies significantly between paediatric and adult oncological institutions, between regions, and between countries. Within the STRONG-AYA ecosystem we acknowledge that there will be some misalignment of data fields collected within an institution across time, and also between institutions at any given time. Datasets collected in the past will not have been structured the same way as newer ones as organisations learned and developed. As a result, across the Consortium, there will be a high heterogeneity in the collection of some domains, particularly those reliant on PROs, and only some will have been recorded systematically. As PROs will, alongside clinical data, drive insights related to tailored treatment and care, STRONG-AYA is committed to standardising the type and method of PRO data collected across the Consortium for the purposes of interoperability and sustainability. Within STRONG-AYA we will strive for a standardisation of an optimum minimal level of data. To ensure the data ecosystem is used as effectively and efficiently for analyses and insights, the Consortium relies on the work in WP1 related to the COS, WP2 in relation to data minimisation, WP3 in relation to joint controllership agreements related to data and visualisation methods, and importantly local ethical reviews for the appropriate collection and use of local data. Going forward, by establishing a COS for AYA care and using data-driven insights gathered via our pan-European ecosystem, STRONG-AYA will share insights with policy-makers to address and reduce inequalities across the entire disease pathway of different cancers in AYAs in different regions in Europe, towards our aim of establishing a European-wide standard of care. STRONG-AYA will enable AYA healthcare quality outcomes monitoring by tracking access, such as by AYA nurse specialists, fertility consultants, and social workers for discussion about work and establish benchmarking between the different countries on key performance indicators. 11.2.6. Inclusion, focused upon patient benefit STRONG-AYA tools and activities are patient-centred and based on the European Ethical Principles for Digital Health, following the principle of ‘nothing about me without me’. These ethical principles, centred on the interests and values of patients, are the foundation the Consortium builds upon, are reiterated and reflected in the development of the FL ecosystem, and built through stakeholder engagement and reflections from the Patient Advisory Board. Importantly, any tools and actions within STRONG-AYA are meant to complement and optimise face-to-face healthcare, patients are informed both of the benefits and limits of tools developed by the Consortium, and they can customise their interactions with these tools. Furthermore, through work currently underway in WP4 by engaging patients in the development of the research and tools, STRONG-AYA aims to ensure patient-facing tools and resulting conclusions are accessible. Namely, STRONG-AYA will ensure patients: Are able to easily and reliably access/retrieve their own STRONG-AYA health data (through their own hospital systems) See their own health data in the context of other AYA similar to them
STRONG-AYA – No. 101057482 – D4.7 18 Easily access information on how their health data has been accessed or used and for what purpose, such that they should be able to easily enhance, grant or remove access to their health data depending upon their perspective on each ‘case of use’ of their data, and exercise this right freely Any tools developed by STRONG-AYA need to be inclusive and accessible to all (including by people with disabilities or low levels of digital literacy). They should be intuitive and easy to use, patients should have access to training to help them understand digital health tools, which in turn should be implemented in care routines that also include human communication. Patient representation on the CfC will ensure that a patient voice is embedded in the ecosystem’s governing structure Clinical trials and medical research, historically, have been driven by the questions and concerns clinicians and researchers wished to investigate. However, it is patients who are directly impacted by the outcomes of research. Their questions and concerns should be placed on equal footing with clinical research areas of interest. By involving patients in the process of co-creation, STRONG-AYA will create a platform that serves patients and empowers them by further increasing their health literacy. Following upon the above, STRONG-AYA will equip AYAs with cancer with the data to better optimise their healthcare, and support shared and personalised decision-making between AYA and health care providers. This also extends to the control of, access to, and use of AYA’s data in the context of research and healthcare. Importantly, by involving patient groups in the conversations around the actions and operations within the ecosystem and the co-creation of the patient platform and methods of data visualisation, the Consortium is committed to ‘meaningful’ patient participation leading to actionable patient-centred and patient-informed insights. The effort of co-creation is expected to result both via the local/national-level governance and ethical systems which are dedicated to incorporating patients’ views, but also through the guiding principles of the consortium. 11.2.7. Privacy, and confidentiality-preserving By adopting a FL and federated analysis approach, the privacy and confidentiality of the participant is protected. Details of how this is performed can be found in D3.3 Technical Blueprint11. Here we simply offer a general overview of how STRONG-AYA will meet this foundational principle. Specifically, we will adhere to both GDPR at an international level and local data protection regulations in what concerns the levels of data anonymization or pseudonymisation for federated learning. Within STRONG-AYA we make a distinction between: 1. federated analyses that do not exchange ANY individual patient information, ONLY group-level statistical summaries (e.g. averages, confidence intervals, regression coefficients) and 2. analyses that during the computation process will unavoidably exchange outcome labels of individuals (e.g. alive/deceased), but these would need additional data in order to identify the individual. The first point encompasses one of the technological cornerstones of the project - federated analytics. Statistical analyses and descriptive statistical visualisations on groups of participants will be performed without transmitting any individual patient-level data outside of the data owning institution. The second point relates to STRONG-AYA supporting federated learning – a method of developing epidemiological statistical models on a group of participants, without transmitting identifiable patient-level data between institutions. Some epidemiological models will require the exchange of the frequencies of certain outcomes to answer research questions – such as ‘alive (or deceased)’, ‘diagnosis positive (or negative) or a time interval (such as ’90 days since a certain event’). In some circumstances, such an outcome may be
STRONG-AYA – No. 101057482 – D4.7 19 characteristic to only 1 participant. However, without individual additional data the individual cannot be re-identified, but such specific individual data will not be accessible via the ecosystem and federated learning infrastructure. Therefore, privacy protection in federated learning and analytics is based on the unfeasibility of reverse-engineering a particular participants’ identity on the basis of summaries, statistical models or other such combined data analyses results based on groups of participants (which are also typically presented in reports and publications)11. Finally, the development of the COS is based on the idea of data minimization – structured data for the COS is extracted by investigators within each institution and placed in their respective STRONG-AYA secure environment controlled by their own institution and governed by their own data protection regulations. All data placed in these areas will be checked to ensure it is pseudonymised or anonymised (as necessary) consistent with the STRONG-AYA Data Protection Agreement to avoid the re-identification of individuals and in accordance both with GDPR and local regulations. For instance, identifiable information which is not necessary for analyses such as birth dates and postcodes is not allowed. No dates will be stored in the STRONG-AYA repository, only time ranges. Similarly, to avoid re-identification, no analyses will be run if the count of participants at the intersection of the criteria query falls below 10. 11.2.8. Scalability and sustainability As outlined in the Business architecture (D2.119), the STRONG-AYA Consortium is committed to the principle of scalability of its resulting activities, tools, and pan-European ecosystem – namely, translatable to other contexts - and their sustainability – namely, to last beyond the lifetime of the research project. This will be done through the operations of the ecosystem outlined here, through the delivery of the COS, through federated learning mechanisms, the delivery of usable and inclusive tools, and adaptability to changing ethics and governance landscapes. The development the COS for AYAs with cancer enables STRONG-AYA to offer new opportunities for the collection of homogeneous prospective data, which will support new research and healthcare innovation. The agile and adaptable approach of FL allows flexibility in incorporating existing systems and practices in larger ecosystems, allowing for the accommodation of new local contexts and ecosystems. Tools which will enable and motivate stakeholders to join and work together are the websites, specific ways of set-up of local databases, portals for stakeholders to enter and access desired information, and finally the access to the expertise required to operate ecosystems at local and European levels. Finally, STRONG-AYA will remain responsive to emerging and unanticipated ethical issues by adhering to iteratively developed ethical guidance throughout the duration of the project, including patient-centered and patient-informed ethical guidance. The sustainability of these aspects will be ensured via the CfC. Continuous ethical monitoring is important because not all ethical issues can be foreseen – some concerns and implications of building a new research data infrastructure and configuring new alliances will have unanticipated consequences. By using an iterative ethical guidance method, STRONG-AYA aims to respond to relevant ethical and governance issues with a reflexive approach, attending to concerns as they arise from the developing research infrastructure and collaborative research actions. To do so, STRONGAYA draws on expertise within the Consortium to organize suitable platforms and formats of patient 11D3.3_TECHNICALBLUEPRINT_STRONGAYA_v2.pdf 19D2.1_ECOSYSTEM_BUSINESS_ARCHITECTURE_STRONGAYA_2023_FINAL1.pdf
STRONG-AYA – No. 101057482 – D4.7 20 consultation that result in meaningful patient participation and leads to actionable patientcentered and patient-informed insights20. Through these, STRONG-AYA aims to generate a strong culture of successful and sustainable means of working together as well as the integration between clinical epidemiological and data science processes with the views of stakeholders. There is an expectation that the ecosystem will speed up analyses on cancer in AYA which should lead to faster decision-making in terms of treatment and care. 11.3. Aligned incentives and motives, across Consortium members All Consortium members are expected to be transparent in communicating their incentives and motives for joining a FL Consortium. This will help the functioning of the pan-European ecosystem by offering a general view of the expectations and responsibilities of different users at local, national, and ultimately pan-European levels. The STRONG-AYA Consortium is committed to satisfying the goals of each partner, flexibly, hence transparency and specificity in incentives is needed for a purposeful collaboration. The alignment of incentives and motives relies on a keen understanding of each institution’s capacity to achieve its goals through the joining of a FL ecosystem. This translates into having the means and mature enough processes for data collection and storage, or processes which can be optimised with some support - for instance ensuring alignment with the technical requirements outlined in the Technical Blueprint (D3.3)11. This also pertains to the description of goals, objectives and key performance indicators leading to these (see D5.2)21. This will mean an alignment (with support from the Consortium expertise and experience) on capacity for clinical leadership, administrative power, scientific, technical, and ethical governance knowledge. The ongoing incentives and motives for the existence of the STRONG-AYA Consortium and pan-European ecosystem are reflected in the principles outlined above. In short, they can be summarised as (for details, see D2.2 Code of Conduct18): 1. Improvement in the provision of healthcare services: The development of value-based clinical practice, that becomes responsive to new knowledge in real time Improving and accelerating the implementation of real-world insights into clinical practice. 2. Performing ethical research on large datasets: Encouraging international collaboration for research and clinical expertise Improving and expanding on the applications of available data in research Understanding and leveraging institutional capacity. 3. Improving the quality of patient outcomes: Defining and identifying patient needs – for research and clinical purposes Improving the reach of research and understanding of patient outcomes for patients from diverse backgrounds Improving patient outcomes through the use of real world data. It is expected that different members will have differing priorities among these incentives for wanting to participate in the Consortium and this is acceptable. Furthermore, it is acknowledged that institutional incentives and motivations to join a Consortium can go beyond patient benefit in terms of the queries they 11D3.3_TECHNICALBLUEPRINT_STRONGAYA_v2.pdf 18 D2.2_ETHICALGUIDELINES_CODEOFCONDUCT_STRONGAYA_v1.pdf 20 D4.6 _OPERATIONALPLANSNATIONALECOSYSTEMS_STRONGAYA_v.1.pdf 21 D5.2_KPISandSUCCESSMETRICS_STRONGAYA_v.1.pdf
STRONG-AYA – No. 101057482 – D4.7 21 may want to run. These can range from discovery (i.e. increasing volume of data to help manage rare diseases), improving and expanding on implementations of past initiatives (i.e. sharing workload on diverse datasets), or encouraging sustainable international collaboration (i.e. tangible outcomes can foster collaborative practices). Whichever the motives of a potential new member, the main characteristic needs to be clinical actionability related to decisions on what queries can be run on the federated data. This also applies to how data is collected and monitored, what analyses can be pursued and which analyses can be shared or not with patients or healthcare providers. Hence, the purpose of data being collected, contributed to the data ecosystem and analysed will have a clearly defined clinical purpose for patient benefit, in line with the definition of ‘bona fide’ research as outlined in the Code of Conduct (D2.2) and regulated by local ethics boards. This will be a requirement for all professional stakeholders (i.e. new members of the Consortium) including healthcare, academic institutions, charities, policy makers, industry partners, etc. The incentives guide the type of outcomes which are acceptable to our ecosystem, which outcomes are monitored, and therefore the design of use cases and data queries, how the local ecosystems are organised and the organisational capacity allocated or due to be developed for an active membership in the ecosystem. The CfC and existing Consortium members will promote the alignment of the different national ecosystems with these incentives and overall mission of the project for the promotion of scalability and sustainability. 11.4. Managing national and institutional differences Within the definition of the governance model with its structure, principles and alignment of incentives and motives, the Consortium acknowledges the differences between institutions and countries in their capacity, capabilities, and limitations in what data can or cannot used for federated learning or analyses. There is an expectation that not all countries and institutions will be able to fully meet all requirements (technical, administrative etc) and there will be gaps in policies. The ethos of a collaborative pan-European ecosystem is that these gaps can be bridged as part of the working practices. A task within the ecosystem needs to be the transparent communication and discussion of these potential differences and gaps to allow room for their troubleshooting and secondary plans to achieve explicit goals. Before joining the Consortium, there is an expectation of local discussions with leadership, technical and legal teams akin to those pursued in D4.4. and D4.6 to identify local capacity. Local evaluation of data availability and metadata structuring, to identify strengths and opportunities for development at each national ecosystem level may be necessary. These interviews can be had with a third, impartial organisation but the resulting insights need to be shared with the CfC as part of the new member application process. Internal interviews and audits for insights into potential knowledge gaps should prompt local discussions (as exemplified by D4.4 and D4.6) and should incorporate at minimum the following categories: A. Provision and availability of retrospective data: Is data available and can it be analysed? *The questions below apply to retrospective data. Approvals in line with GDPR and local legal/ethical guidelines are expected to be arranged by consortium members. What data is currently available, from previous studies or clinical practice? Did patients consent to their data being processed for public health research? Can the existing data be FAIR-ified? What is the local complaints procedure (i.e. immediacy of action, timeframes for reply etc)? B. Data collection and consent norms for prospective data: How is data collected?
STRONG-AYA – No. 101057482 – D4.7 22 *The questions below apply to prospective studies. Appropriate Ethics Board approvals in line with GDPR and local legal and ethical guidelines will be expected to be in place in each location. Do patients know why the data is being collected? Do patients agree to share the data? Do patients know why the data is being shared? Do patients understand how it is being shared? What is the local institutional process of communicating results back to patients (if necessary)? Does data collection follow pre-existing regulations and rules? What is the local complaints procedure (i.e. immediacy of action, timeframes for reply etc)? C. Operational norms and standards: How does each institution operationalise the Consortium internally? *Please read the current General Consortium Agreement for an example of how these questions can be answered by an applying member. From whom will you accept a data query? For what purpose will you use (or allow use) of the data? What results will you allow to be returned? Who adjudicates disputes or ambiguous situations locally, in regards to data collection and analyses (either from participants or researchers)? What is the threshold to entry for a new member of the federations? D. Technical standards: What are the technical standards (i.e. How are the datasets managed to guarantee data security, data integrity, and patient privacy)? *Please read the current D3.3. Technical Blueprint and FAIR principle above, for an example of how these questions can be answered by an applying member. Within STRONG-AYA these are managed within WP3 with suitable software with clear sustainability plans available from IKNL. How do you ensure security of queries within the transfer process? How is patient privacy protected? Do you make your build standards transparent? What is the change management process? What level of interoperability needs to be reached (in terms of software, APIs, and datasets)? What are your mechanisms for ensuring data integrity? How is data standardisation currently achieved? E. Legal and ethical governance standards: What are the legal bases for consenting, collecting, and using patient data in research in each country and institution? Does a joining institution have consent and research process protocols and practices in place? Do these allow (and to what extent) for insight sharing as part of a federated learning ecosystem? As an example, we expect at minimum to identify the following list of potential (not exhaustive) differences and gaps, which can then be leveraged and managed within the Consortium: 1. Availability of retrospective (historical) clinical and PROs data within centres 2. Ethical and governance rules towards the re-use of previously collected data for additional analyses or insights
STRONG-AYA – No. 101057482 – D4.7 23 3. Limitations in patient information and consent procedures at national levels for retrospective and prospective studies 4. National restrictions in collecting specific prospective data (e.g. in France collection of data pertaining to ethnicity is not permitted) 5. Ethical limitations in conducting particular analyses (i.e. cancer incidence by ethnicity) 6. Gaps in technical or administrative expertise in handling large data repositories 7. Potential differences or gaps in implementing software solutions, APIs, and managing data system upgrades or improvements in technical components 8. Time limitations in the storage and use of existing or new data. 12. Operations of the federated learning pan-European ecosystem 12.1. Summary As described in section 5.2 and particularly in Figure 1, our pan-European ecosystem relies on the appropriate functioning of its parts – each local, and over time national, ecosystem with its own operational environment (described in detail in D4.4 and D4.6). Several interviews and focus groups with representatives of each national ecosystem evaluated the local and national ‘Preparedness’ of the human resources, technological, and data management and security infrastructures. The Operational plan below acknowledges that our pan-European ecosystem can only evolve as fast as its parts (see Figure 1); we incorporate knowledge gathered through individual national interviews related to the healthcare and technological system readiness for a cross-country (and later cross-European) development of an ecosystem. While this may not be possible at a national level in each locale as yet, in certain parts of Europe our initiative has opened the doors to the development of the appropriate infrastructure to enable this beyond a local level in each nation. Our initiative highlighted the current threats and opportunities towards the goal of creating national ecosystems tharcan then be integrated with an international ecosystem. The status of ecosystem preparedness in each locale to develop towards a national and then European ecosystem is described in D4.4 and D4.6. D4.4 offers a detailed plan of the actions and issues all institutions within the STRONG-AYA consortium need to fulfill to operate their local ecosystems and build national AYA cancer ecosystems, including descriptions of people/human resources, priorities of actions, and pragmatics of implementing the federated ecosystem. D4.6 set out of number of individual actions from each partner that need to be met to operate their own local and national, and therefore underpin our pan-European ecosystem. These actions formed the individual roadmaps for each centre whose data will contribute to the European federated data. The update on national system preparedness is available in D4.9, which includes an update on system preparedness one year later following the initial evaluations described in D4.4. Here we summarize the actions still needed across nations to reach the ambitious goal of a pan-European eco-system with multiple stakeholders at various levels of preparedness. Our plan below outlines the development and integration of national ecosystems into a pan-European data-sharing and research-driven ecosystem for AYAs with cancer. It includes critical actions around human resources, technology, data management, and security, and aligns with the overarching principles of data federation, ethical compliance, and stakeholder collaboration, ensuring sustainable growth across national and European frameworks.
STRONG-AYA – No. 101057482 – D4.7 24 12.2. Operational objectives Establish a collaborative data ecosystem for AYAs with cancer, including in each participating location and across nations. Ensure secure, scalable data management and monitoring across local and national ecosystems. Drive patient-centered care and outcome-based analysis through the integration of COS and clinical data associated with the COS at local and national levels. Promote sustainable collaboration across all participating nations. 12.3. Scope of Work Support the development of individual local and national ecosystems, each tailored to local contexts, regulations, and healthcare needs. Implement a federated data infrastructure that allows secure sharing across local, national, and international ecosystems. Establish governance, data security, and ethical frameworks consistent with national and EU standards. Define local actions, milestones, and resources needed for a functional ecosystem. 12.4. Work Structure: 1. Ecosystem Setup Define national ecosystem leaders and teams. Establish local infrastructure for data collection and storage (e.g., Secure Research Environments). 2. Technology and Data Integration Implement federated data systems for secure cross-border data access. Install portals for data submission and monitoring in each ecosystem. 3. Governance and Compliance Align local ecosystems with GDPR and relevant ethical frameworks. Obtain legal approvals and consent forms for data sharing. 4. Stakeholder Involvement Engage patients, healthcare providers, and local governments. Implement cultural and language adaptations in patient portals. 5. Continuous Monitoring and Control Regular security audits. Ongoing risk management to address potential cyber-attacks and resource constraints. 12.5. National Ecosystems The status of this work as of 30/09/2024 is detailed in D4.4 and 4.6, following detailed interviews with ecosystem leaders in each location and nation.
STRONG-AYA – No. 101057482 – D4.7 25 Stakeholder engagement in collaborative communities is strong in each nation, except for national data policy-makers. Data access, currently based upon data altruism and mutual trust, is delayed, due to some academic delays (in defining the Core Outcome dataset to be collected), information technology system limitations and upgrades, and research culture barriers for AYA research, each in some locations. Data management expertise and capacity is in place or imminent in each partner centre, to ensure metadata and data quality. Recent research grant awards are facilitating widening our data access in France. Other research grant applications can contribute to the breadth of our pan-European ecosystem, if funded. There are successes and ongoing progress plans in each nation The Netherlands: Secure data storage is in place in local hospitals covering the whole nation. Data transfer agreements and protocols for existing data from existing studies (e.g. COMPRAYA) are ongoing. UK: Aspirations to obtain existing national data from NCRAS have been slowed by reduced governmental funding. Clinical leaders for AYA cancer in NCRAS are addressing this to achieve a specific limited national dataset, through key stakeholders (TYAC, Christie NHS Trust). France: New research grant aware will enable wider data access. Regulatory approval for COS data collection is delayed by the finalization of the COS. Italy: Local research dataset in place, from previous projects. A secure infrastructure for data is in development. Poland: Stakeholders in Warsaw are building a local research environment. 12.6. Technology and Data Management Federated Data Infrastructure: Each nation retains local control over its data, contributing aggregate data summaries for cross-European analysis. Data Governance: Ensure that the data collection and transfer adhere to the FAIR principles (Findable, Accessible, Interoperable, and Reusable). Tools: Web-based applications and portals will allow real-time and on-demand analysis of the collected data. These tools will integrate patient and provider feedback loops to enhance decisionmaking. 12.7. Human Resources Management Each center is recruiting and training local data managers and support staff to ensure compliance with data collection and analysis protocols. Organize workshops to train local healthcare providers on data collection and interpretation for AYAs. 12.8. Data collection (COS and clinical information) Monitoring and Uptake Implement training for both clinicians and patients to increase the usage of COS. Regular evaluations of COS integration into routine clinical practices in countries running the implementation study.
STRONG-AYA – No. 101057482 – D4.7 32 concerns within the Consortium to be reported and resolved in a timely manner with appropriate replies to stakeholders. Partners must agree upon a balance between data utilization, data security, and preferences of the data providers. However, work to make the data analyses useful to others beyond the initial Consortium is vital to sustainability, so this needs to be subject to further work across the STRONGAYA ecosystem as it grows and this document will be updated accordingly. In order to ensure the sustainability of the ecosystem infrastructure, through user access management and monitoring, Consortium members will be required to contribute and collaborate on ongoing funding initiatives for the ongoing use and maintenance of the system. While currently access to these services is free of charge, in the future, as the consortium grows and expands access to the eco-system infrastructure may develop into a subscription-type model. 13.5. Resolution of disputes and ambiguous situations Disputes, complaints or ambiguous situations at local levels are expected to be resolved using the policies and processes specific to each institution as identified through PLUTO. At the Consortium level, STRONG-AYA, where reasonable, will employ a bottom-up approach to adjudicate disputes or ambiguous situations (Figure 3). These may relate to decisions around the joining and membership to the data ecosystem and Consortium, concerns related to a new use case or research methods from any stakeholder (professional or patient). Any disputes arising between STRONG-AYA Consortium partners should be solved amicably and respectfully. Ambiguities on any given criteria to establish thresholds of access or how rules and standards of operation are applied, will be discussed at the level of specific WP with the help of the work package leads and the Project Coordinator. If unresolved, the Project Coordinator will escalate the issue to the CfC (specifically the Coordinator and the Scientific Coordinator) who will use mediation, expert and referent powers to objectively aim to resolve the issue. If the ambiguous situations or disputes cannot be resolved at this level, the Coordinator may appoint a dispute resolution panel or opt to refer the issue to the General Assembly for wider consultation if needed. If a partner wishes to complain about other partners in the Consortium, the complaint should be documented in writing. Such a complaint will be addressed to the CfC (Coordinator and the Scientific Coordinator) and the Project Coordinator in the first instance. If a complaint resolution panel cannot resolve the issue, the General Assembly will be convened and will vote on a resolution to reach a binding solution. Figure 4: Process of dispute and complaint resolution within the STRONG-AYA Consortium
STRONG-AYA – No. 101057482 – D4.7 33 13.6. Submission of use cases The STRONG-AYA Consortium has developed an initial priority list of how the data within the ecosystem infrastructure can be used by stakeholders for research and education purposes. This priority list can be seen as ‘research questions’ or ‘database queries’ or, as they will be referred to henceforth, ‘use cases’. These can only be submitted by Consortium members with full or partial user level access (see Figure 4 for a depiction of the process). Figure 5: Process of submitting and approving new Use cases by Consortium members The current list of use cases has been developed as a list of key questions, which could be answered by analysing data within each institution, aligned to the motives, goals, and performance indicators of the STRONG-AYA pan-European ecosystem as a whole and those of each of its members (local and national ecosystems). Consistent with the principles of standardisation and patient benefit, the use cases are aligned to the COS defined in WP1 (see D1.2) and they have been further defined and prioritised in collaboration with service users (patients and carers), healthcare professionals, researchers, health service managers, and policy makers (see D4.8). It is envisaged that this list of use cases is dynamic and will evolve over time. Consequently, here we set the rules for the development and submission of new use cases to the Consortium. An approved member of the Consortium can choose to submit a use case to be reviewed by the CfC (for an example of how the initial use cases have been defined, see D4.8). The criteria these will be evaluated against are: Demonstrable clinical benefit: the use cases need to lead directly or imminently (within maximum 5 years) to improvement in clinical or psychosocial care. Demonstrable co-design with relevant stakeholders: members need to describe the process through which the use cases were identified, formulated, and prioritised with all stakeholders involved (patient groups, leads or members of other work packages, researchers, policy makers etc.). Description of legal and ethical considerations: use cases submitted for review should include an assessment of legal and ethical risks and mitigations procedures from local ethics/governance bodies and consistent with GDPR. Description of administrative, technical, research, and clinical capability: once a use case is submitted and approved it will produce an output which can be used by the member(s) of the Consortium. The implications of a given insight need to be assessed and reassurance offered that
STRONG-AYA – No. 101057482 – D4.7 34 the Consortium member(s) have the capacity to use the insight for patient benefit and prevent any unintended harm. Plans for growth: as part of the process of submitting a use case for review, Consortium members need to demonstrate the practical steps of how the use case will lead logically and in a timely manner to patient benefit, for themselves and/or other Consortium members, with associated KPIs. 14. Rules on using the ecosystem for federated learning and analytics Once a Consortium partner has defined their local governance model, set up or upgraded their capabilities (technical, administrative, etc) to meet the standards of the Consortium, have had their use cases approved, and have implemented the API, they are ready to use the ecosystem for FL and analytics. This action is led by its own specifications and rules defined below. 14.1. Active contribution to the data ecosystem New and existing partners are expected to contribute to the data ecosystem with their own datasets, which will support the sustainability and continuous development of the ecosystem. Honoring the principle of Standardisation defined earlier means that partners will need to identify ways to improve the structuring of their existing data and implement these improvements in new data collected. To achieve this, partners and their local investigators may want to use metadata (information about data such as its origin, meaning, measuring scales, location, timing of collection, limits on storage duration, etc.) which will support the use of algorithms for federated analytics. While the STRONG-AYA commits to offering the tools and data science support to develop this, the uptake of local investigators is crucial to the understanding of data coding and labelling subtleties. There is an expectation of differences in the labelling and formatting of datasets between local and national ecosystems. The Consortium welcomes differing data structures, as long as it can be joined up by metadata making it possible to pursue federated learning and analytics across the ecosystem. To assess internally whether active contribution is possible, each Consortium member needs to self-evaluate whether they can meet certain expectations (for more details, see D3.1. Data management plan and D3.3. Technical Blueprint), namely: Each Consortium member is expected to have a range of existing data already collected or being collected, covering PROs, clinical and service use data with relevance to AYAs with cancer. Capacity to provide metadata on these is necessary. The existing data will be associated with their own data collection workflows and data management practices characteristic to each institution, independent from STRONG-AYA, which can be transparently described. It can be expected that any new joining partners as well as existing ones will have a suitable local electronic patient record and a derived clinical database. Each centre has an existing (or at minimum in advanced development) database held in a secure research environment suitable for special category STRONG-AYA data. This will be used for the hosting of data locally, running local computations, and for the storage of partial summaries to integrate into the Consortium-wide statistical model as part of the federated learning process; Each joining partner needs to provide the metadata of the data they will be providing for the federated learning ecosystem to ensure and maintain the interpretability and interoperability of the system.
STRONG-AYA – No. 101057482 – D4.7 35 14.2. Abiding by the technical standards The Technical Blueprint (D3.3.) details the technical requirements, the plan to leverage existing expertise, and then provide support to extend and improve local systems. The Consortium will provide support for partners to follow standards for the management of datasets to ensure data security, integrity and privacy preservation. STRONG-AYA promotes, encourages and empowers autonomy, flexibility and sustainability at the partner level. Each institution’s contribution to the COS dataset in STRONG-AYA will remain fully controlled, and managed locally, by the leading investigator representing that institution. The technical teams of each partner institution in the Consortium must strive to improve their own datasets and their management with agreed, consistent system upgrades. The system needs to be able to use the APIs for federating data which will be programmed centrally (following use case approval) by members of the CfC following the principles and standards of the Consortium’s governance model (Section 3). As the goals and objectives of the Consortium grow or additional partners join the Consortium, changes will need to be made and a clear change-management process will be further outlined within this dynamic report. 14.3. Privacy preservation The STRONG-AYA Consortium is dedicated to the concept of ‘privacy-by-design’. No data will be exported locally from where it has permissions to be held. Only authorised staff of a given institution are allowed access to the line-by-line data inside their data station. No one outside of that institution will be given access to this. Any remote access to the information (anonymised or pseudonymised) stored in a specific institution is completely forbidden. However, each partner will be able to see the output of analyses that includes the data they have contributed, if they had been granted access to the layer of the ecosystem website holding those outputs. For example, patients may be able to see their own personal line-by-line Strong-AYA data (through a portal into their local Electronic Health Care records), but not the line-by-line data of others. However, they may be able to see comparative summaries of their own data against summaries of similar data of other similar patients (e.g. they may be able to see their reported PROs data versus average responses of other people with the same diagnosis). Data relevant to the federated system will not be accessed by third parties without explicit, unanimous approval by all members of the Consortium and without explicit consent from patients contributing data to the Consortium. Protection from data breaches is ensured through a set of robust penetration tests – a report is available from the CfC upon request. Protection from data leakages is ensured through the local cybersecurity standards each institution has to follow, commensurate with the requirements set out in the Technical Blueprint and Data management plan. STRONG-AYA encourages partners to ensure the availability of a compensation fund, in the unlikely event they become a victim of cybersecurity attacks. Data will not be used for the purposes of identifying patients or attempting to identify patients in any circumstances. Failure to follow this rule will result in immediate and permanent removal from access to the Consortium. Any data breaches will be investigated using local data protection processes and by the CfC commensurate with the Data protection agreement. 14.4. Supporting interpretability and interoperability A large part of work pursued by the STRONG-AYA pan-European ecosystem focused on ensuring the interpretability and interoperability of data at cross-national level through adherence to FAIR principles. This
STRONG-AYA – No. 101057482 – D4.7 36 was done through the development of the WP1 COS and use of metadata. The COS is an example of alignment of the data collected under the same definitions in various nations and institutions. This way, a common language is ensured as standardised, more homogeneous information can be collected and provided at local levels, which insures interoperability – any partner will be able to use the data collected with a clear knowledge of what it represents. The use of metadata locally and across the consortium also ensures the interoperability of the data infrastructures. For example, haemoglobin may be measured across all participating countries. Metadata related to this data point in each institution will contain information of the measuring units in which it is represented (e.g. grams per decilitre of blood or grams per litre). While all these units of measurement are eligible and acceptable to be shared without a need for additional transformations, metadata associated with the variable is necessary to ensure algorithms (APIs) can be programmed to run transformations automatically and that the output can be easily interpreted by the user. Another example is that of specific interpretations of less known variables – such as nicotine exposure. This can be encoded as ‘smoking cigarettes’ or ‘chewing tobacco’ as lower level sub-classes of nicotine exposure but it can also be accompanied by quantity specifications such as number of pack/years of consumption. Similarly, one may define a higher level class of ‘T4’ for tumour staging, with T4a and T4b as sub-classes. In this case, searching for ‘T4’ will return all participants with T4, T4a, and T4b, whereas searching for T4a will only return these specific participants with a T4a. The programming within the API will be based on the embedded hierarchical relationships between variables. To enhance the re-usability of local datasets the federated queries (developed on the basis of use cases) will convert specific data labels provided via metadata into generic (i.e. schema-independent RDF objects) and then adding study-specific annotations. Thus, consistency in the metadata shared by each partner is necessary to ensure the system is interoperable and a correct interpretation of the COS data and insights. While the COS has already been defined, there is some flexibility whereby new outcomes of interest may be added if associated with a clear rationale and provided their measurement does not raise ethical concerns such as discrimination. These aspects will be evaluated by CfC in collaboration with WP2. In the event that a prospective partner is only collecting a specific, niche type of data without plans of expanding this in volume or scope, this will influence their ability to make the Consortium successful. Consequently, when the type of existing or planned to be collected data is limited in availability (e.g. very rare tumour types) partners will need to provide information on future data collection plans and/or growth trajectories for data sets as a way of helping themselves and CfC determine institutional capacity. The Consortium appreciates the time and effort necessary for these exercises if this work is not already in place as part of the local data management system, and is committed to providing support when necessary and feasible. Differing precise types of data should not be a roadblock to a successful data ecosystem and Consortium, but there needs to be a recognition that the absence of such information may limit the longterm operation of the Consortium and will limit the ability to derive new insights via federated learning and ultimately deliver on shared goals and incentives. 14.5. Tracking success with KPIs As part of the application process, potential new members will need to define and provide a breakdown of key performance indicators (KPIs) associated with their incentives and goals (see D5.2 as an example). This is also a requirement for ongoing members, hence KPIs will be characteristic of each new project a partner develops and will related directly to the existing and new use cases submitted for CfC review. The data insights accessed via the federated system and subsequent clinical or research findings will be tracked in line
STRONG-AYA – No. 101057482 – D4.7 37 with agreed KPIs, which will be dependent on the build of the API (i.e. the algorithms sent through the federated ecosystem to produce summary statistics). 14.6. Using the infrastructure For the use of the STRONG-AYA infrastructure members need to: Abide by a purpose that is within the list of acceptable authorized purposes for use (e.g. education or research) and consistent with the limitations of local registers. Any purpose not approved initially should be discussed and approved with the Consortium and data providers. A disclaimer will be available on all visible part of the STRONG-AYA infrastructure stating that the available tools should be used only for research or education purposes. Concrete examples will be made available. Follow the recommendations and training provided on how we are legally permitted and have agreement to use the system for the different purposes such for patient information, research, etc. Users will be made aware of the limitations of using AI solution and the challenges to incorporate AI in the patient-doctor interactions. All Consortium members need to raise user awareness about the risk of overreliance on the system. Users should take the time to understand and to challenge the outcomes from the system before implementing them especially in clinical setting. Clinicians should make first their own diagnosis and then compare it with the output from the system. Training materials, disclaimers, and output descriptions on the STRONG-AYA web interface will be adapted to the type of users – healthcare professionals or patients – in conjunction with other deliverables. Below (Table 2) we identified a list of potential risks associated with the use of a special category data ecosystem and some of the ways in which these risks will be mitigated by STRONG-AYA. This list is not exhaustive and additional risks will be added as they are identified. Importantly, while some the levels of risks will be different for patients versus professional users, here we offer a generic overview. For additional information please refer to the Code of Conduct (D2.2) and Grant agreement. Risks Mitigation rules Risks related to patient participants Patient re-identification Many people li ving with a rare disease may be easier to re-identify due to the low number of other individuals in the world with the same variants of a specific rare disease. In the case of extreme rare events reidentification can take place through the crossing of data from the same patients from different but accessible sources. Attempting data queries of rare events with minor changes in settings/rules of data searches can lead to the attributes of participants being inferred. There will be an ongoing effort to seek data to achi eve greater volumes of diagnoses for people with rare diseases. The software automatically excludes any computations on 10 or fewer subjects in any given data station. No analyses will be permitted internationally if data cells fall under the agreed data suppression thresholds (such as 10 cases per cell). These restrictions will apply both for absolute data quantity but also for the amount of new data added
STRONG-AYA – No. 101057482 – D4.7 38 to the ecosystem since the last analysis by any one API. This prevents analyses and re-analyses inadvertently providing identifiable information or participant attributes to be inferred. Risks Mitigation rules Patient withdrawal from study Patients may wish to withdraw their contribution to STRONG-AYA following the pseudonymisation of data. Bein g able to withdraw consent is a right conferred by GDPR. As the initial pseudonymisation of data is done locally by each institution, and the request of withdrawal will be directed at that institution, such cases will be dealt with locally. Consortium members managing local data station are expected to have the means (within the limits of their own governance systems) to manage changes in consent status and re-identify participants, if necessary, by accessing their own data store separate from the pseudonymised dataset used for federated learning. Patient participants do not engage or do not adhere to survey completions Lack of patient engagement and adherence to survey completion is common in PRO research studies. To mitigate this, patient advice will b e sought in the development of tools and outcomes to be collected through the STRONG-AYA Patient Advisory Board. They will ensure materials are visually engaging, sufficiently informative and aligned to patient needs to motivate the completion of the questionnaire and to ensure retention. Personalized feedback is envisaged to motivate patients to continue participation. Guidance and helpdesks will be available. Discrimination or other uses not in the interest of the patients (harmful use) Potential queries in the API may lead to discriminatory conclusions, which could endanger patient safety. The use and integration of AI/big data solution into medical practice brings the challenge of preserving patient rights and interests, while still making full use of innovation. Potential queries of variables within the ecosystem which might lead to biased discriminatory conclusions will be limited through the design of the system. However, it is difficult to envisage all the ways in which a query might lead to discrimination without limiting the use of psychosocial variables. Users of the system will be able to flag particular queries or variables as potentially discriminatory and offer a description of this concern. This will enable STRONGAYA to improve the system over time. Users of the STRONG-AYA research infrastructure should employ it only for the benefit of the patient (on the basis of bona fide research for the improvement of patient survival and quality of life). This is ensured through the tracking of API usage and
STRONG-AYA – No. 101057482 – D4.7 39 its alignment with pre - defined KPIs, use cases, and goals. Users should not use the infrastructure for potentially harmful use that could lead to any type of discrimination or endanger patient’s wellbeing. Any such actions will lead to the removal of the partner from the Consortium. Risks Mitigation rules Distressing conclusions Data generated by the STRONG-AYA infrastructure may require strength in the patientdoctor relationship, such as different cancer outcomes between nations and centres. The infrastructure could provide information to a health care professional or a patient which requires discussion and explanaiton to manage any distress. Appropriate training and guidance will be provided to all users of the system for how insights generated should be interpreted or be communicated to a patient (if not the user). The STRONG-AYA infrastructure will not be a route to specific clinical advice. However, for those who may have particular questions, healthcare professional contact information will be made available for patients to discuss potentially distressing findings. Technical feedback can be provided by IKNL and UNIMAAS as members of a CfC. The interpretation of analyses completed by the APIs will be joined by their purpose, appropriate FAQs, information on the limitations or common misconceptions about the findings. Appropriate disclaimers will be put in place where resulting insights might be distressing to a user, especially in the case of patient/carers. However, a bespoke set of training, output descriptions and disclaimers will be designed for professional versus patient users. Risks to professional members of the Consortium Differences in legal, ethical, or technical standards in local ecosystems of Consortium members Difficulties in creating the legal entity and in developing a viable operating plan for the pan-EU ecosystem due to lack of alignment of national ecosystems operating models. This will be mitigated according to the section on Managing institutional gaps and differences. There will be an expectation of alignment to local governance/ethical standards and processes as well as GDPR from each consortium member.
STRONG-AYA – No. 101057482 – D4.7 40 Risks Mitigation Rules System overreliance for clinical decision making It may be difficult to prevent clin icians from using the system to support their clinical decision-making. This introduces is a risk of overreliance on the system. All final clinical decisions will remain with the clinician, but they may be influenced by the outputs from the system (e.g. quality of life data being used in clinical practice). The STRONG-AYA research infrastructure is devoted to research and educational uses and this rule needs to be followed by all partners. The Consortium will assess, on a case-by-case basis, how the results could be used for clinical decision-making (within the portals to STRONG-AYA and data in local Electronic Health Records) and which limitations are therefore applicable. The purpose of making a specific query in the API will need to be declared by each user at point of use. Appropriate training and guidance will be provided to new users of the system and a feedback procedure from users be made available so that users can challenge outputs if needed. Outputs will include appropriate descriptions and disclaimers commensurate with the type of user that may view them. Users will also be able to contact the STRONG-AYA helpdesk or the local investigators for additional questions and clarifications. Queries outside scope This relates to queries, algorithms, or use cases placed within the API for processing, but outside of the scope of the Consortium and its goals. Data that is collected and available to be queried will always be within the disease area of focus. New queries and use cases are permitted only after due process, considering ethical, motivation-based and wider scientific conditions. Any query, algorithm, or use case submitted via the API which is not aligned to a pre-defined KPI, goal, and approved by the CfC will return an error message. Inconclusive results The user asks wrongly formulated questions which results in inconclusive findings, which could still be applied in clinical practice, hence putting patient safety at risk. There is need for user education and awareness on how to formulate use cases and how to use variables within the system to answer these. Appropriate training and guidance will be provided to new users of the system and a feedback procedure from users be made available so that users can challenge outputs if needed. Appropriate disclaimers will be available for the various outputs, tailored to the type of users. Training, guidance and disclaimers are developed by Consortium members as part of other deliverables. Failure of the technical system due to external attacks on infrastructure
STRONG-AYA – No. 101057482 – D4.7 41 Given the technical nature of FL ecosystems cyber-attacks on infrastructure are a risk in the form of data poisoning; eavesdropping; model poisoning; decommissioning of the server; decommissioning of the institutional data point; penetration of the central service; privacy reversion. These aspects have been considered in detail by the STRONG-AYA technical leads and safeguards are in place as detailed across multiple documents including ROPA/DPIA, Technical Blueprint, etc. Appropriate and robust system penetration testing has been pursued and the report, which identified minimal such risks, is available upon request. Over and above safeguarding at the umbrella ecosystem level, local investigators must follow institutional access and data control policies (multi-factor authentication, strong passwords, not sharing passwords, not giving access to unwanted persons; people interacting with STRONG AYA must be staff who are reporting their actions to the PI). There are measures in place to restrict access to the central server to the minimum number of necessary personnel. Mitigation is by design and encryption – by exchanging only summaries or models and by encrypting telecommunications it’s unfeasible to re-identify a person. Table 3. Risks and mitigation rules 14.7. Intellectual property guidelines All members of the Consortium agree to share intellectual properly and give credit to the Consortium as a whole. If an institution has a pre-existing internal intellectual properly policy, that policy needs to be communicated if different from the statement above. The local institutional policies of the Consortium member will be applied for any discoveries made using data outside of the Consortium’s dataset (i.e. PRO or clinical data available in an institution that is not made available/shared with the Consortium). 14.8. Sharing of benefits - Communication and dissemination of outputs All parts of the work conducted by the Consortium is expected to be reviewed by the CfC and Project Coordinator. They are expected to refine work (such as the COS, Data management plan or this present report), update it, ensure that there is a clear channel to communicate it to partners (i.e. the project Microsoft Teams shared drive and email communications) and to oversee how this work is implemented. Other communication methods will include internal information exchange in pre-defined project meetings, newsletters, and summaries of newly delivered reports to all partners. All and any outputs connected to the data and information (i.e. including images, new ways of working, policy outreach activities etc.) produced within the STRONG-AYA ecosystem should credit the Consortium as a whole.