scieee AI-readable full text Open interactive document viewer

D3.1 - Interim Data Management Plan for STRONG-AYA

Wee, Leonard; Darlington, Anne-Sophie; Aguado, Maria; Lejeune, Stephane; Stark, Dan; Dekker, Andre; van Beusekom, Bart; Couspel, Norbert; Husson, Olga; Hanebaum, Simone

Abstract

This interim Data Management Plan sets out the present vision for decentralized data management. It outlines the decentralised storage of appropriately pseudonymised data in accordance with the principal of data minimization.

Full text

1 A new, interdisciplinary, multi-stakeholder European network to improve healthcare services, research and outcomes for Adolescents and Young Adults. STRONG-AYA – No. 101057482 – D3.1 Deliverable Report Interim Data Management Plan for STRONG-AYA Deliverable D3.1 Short title : Interim Data Management Plan Due date of deliverable: 01/04/2023 Actual submission date: 31/03/2023 Project: STRONG-AYA Lead Contributor Leonard WEE (Clinical Data Science – University Maastricht) Email [email protected] Other Contributors [WP1] Anne-Sophie Darlington ([email protected].uk) – University Southampton [WP2] Maria Aguado ([email protected]) - EORTC [WP2] Stephane Lejeune ([email protected]) – EORTC [WP2&3] Daniel Stark ([email protected]) – University Leeds [WP3] Andre Dekker ([email protected]) – University Maastricht [WP3] Bart van Beusekom ([email protected] – Netherlands Comprehensive Cancer Centre [WP4] Norbert Couspel ([email protected]) – European Cancer Organization [WP5] Olga Husson ([email protected]) – Netherlands Cancer Institute [WP5] Simone Hanebaum ([email protected]) – Netherlands Cancer Institute Emails See after abovementioned persons, in parentheses. STRONG-AYA – No. 101057482 – D3.1 Due date 01 04 2023 Delivery date 01 04 2023 Deliverable type R Dissemination level SEN Description of Work Version Date Drafting work V0.1 - Interim Data Mgmt Plan V1.0 31/03/2023 Description of Deliverable: An initial Data Management Plan will be developed and updated for the final deployment near the end of the project. The Data Management Plan(s) will reflect the technology and situation-specific needs to be addressed by the principles of data privacy by design and by default. It will define data anonymization standards, minimum safeguards and obligations of the data custodians. In collaboration with WP2 it will define a STRONG-AYA information security policy. 1. Summary (max ½ page) STRONG-AYA is a vision for an interdisciplinary, multi-stakeholder European ecosystem utilising widely geographically dispersed data to improve health services, clinical research and patient-centred outcomes of 15-39 year old persons with cancer. This is a special group of people in healthcare, that experiences agespecific and lifestyle-specific issues when diagnosed with cancer, often with drastically-altered quality of life compared to their cancer-free peers. The STRONG-AYA project faces this challenge with an organizational structure and technological innovations that can bring together health data about this special group of people, and learn actionable insights from this data, all the while protecting individual identities by principle and by design. This interim Data Management Plan sets out the present vision for decentralized data management, and may be expected to change in certain details as the STRONG-AYA consortium develops and matures. However, the core concepts contained herein are not expected to change significantly. This document outlines the decentralised storage of appropriately pseudonymised data in accordance with the principal of data minimization. This data shall be made FAIR, but not open due to sensitive nature of personal, medical and lifestyle information (in spite of pseudonymization). FAIRness of the private data is a prerequisite for computer-controlled Federated Analytics and Federated Learning. Federated here means the dissemination of summary statistics and epidemiological models, without the exchange of individual patient-level data. Lastly, the outline of a secure system infrastructure to assure privacy by principle and STRONG-AYA – No. 101057482 – D3.1 privacy by design is presented. A finalized Data Management Plan is scheduled for delivery near the end of the STRONG-AYA project. 1 Introduction General introduction to STRONG-AYA project STRONG-AYA is a vision for an interdisciplinary, multi-stakeholder European federated ecosystem for data availability to improve health services, clinical research and patient-centred outcomes of adolescents and young adults (AYAs) with cancer (i.e., 15-39 years of age at first cancer diagnosis). AYAs comprise a demographically diverse and clinically heterogenous group across the length and breadth of Europe, due to uniquely age-specific and lifestyle-specific issues, as well as drastically altered quality-oflife compared to their cancer-free peers. The STRONG-AYA consortium believes that innovative and paradigm-breaking ways of accessing data and sharing knowledge across institutional siloes is urgently needed to address healthcare needs among this unique group of people. The key pillars of the project are: (1) consensus development of a Core Outcomes Set i.e. COS; (2) infrastructure for collecting, managing and enhancing AYA data among participating countries; (3) disseminating outcomes and data analysis tools at the pan-European level to improve patterns of care for AYAs. A federated ecosystem - for decentralized data management as well as enabling software technology for stakeholders (e.g., clinicians and patients) to be able to perform federated data analysis - was selected as the intended approach to address data fragmentation, bypass adoption barriers associated with centralized patient data repositories, and to be more easily scalable to other centres after end of current project. STRONG-AYA – No. 101057482 – D3.1 Data management that complies with FAIR (Findable-Accessible-Interoperable-Reusable) data principles is a cornerstone of our technical infrastructure. The FAIR principles will accordingly be applied at partner, national and pan-European levels for data, and subsequently digital artefacts generated (e.g., software, publications, reports) from the project work. A further technological cornerstone of the project shall be federated analytics, i.e. performing statistical analyses and summary visualizations of data but all the data stays locally within the original institution. The emphasis is that individual patient-level data will not leave the data-holding institution. STRONG-AYA supports federated learning such that prognostic and predictive (epidemiological) models can be developed and validated, also without sharing individual patient-level data between institutions. Federated analytics (and federated learning) that use geographically dispersed FAIR data requires secure web connectivity infrastructure that is upgradeable, maintained and monitored against external attacks. Userfacing web interfaces hosting applications approved by the STRONG-AYA consortium will be made available, so distinct users will each have differentiated permissions and get access to analytical/learning tools. Scope of Data Management Plan(s) The Data Management plan outlines (1) in general terms the data used in this project, (2) that data is protected by keeping it within the data-holding institutions, (3) implementation of the FAIR principles to make institutional data mutually inter-operable across the STRONG-AYA consortium and (4) availability of the data after the project ends. Additionally, there shall be a number of end products of STRONG-AYA from a data management perspective, which are relevant for the sustainability after funding ends, and for other similar federated FAIR data ecosystem initiatives such as the European Health Data Spaces. The end products include: 1) An information model (i.e. umbrella protocol), that defines the Core Outcomes Set (COS) to be obtained by partners in the STRONG-AYA community. 2) A FAIR implementation profile that describes how the institution-specific data schema is mapped a posteriori to a universal terminology and taxonomic model of the over-arching COS. 3) Open-source software, engineering know-how, and reference implementations of the aforementioned, including synthetic data which helps future partners set up their own FAIR data repositories according to the information model (1) and FAIR implementation profiles (2). 4) Open-source software repository of federated analytics (FAIR models) and reference implementation frameworks for new consortia (FAIR services) that were used in STRONG-AYA so that other users can apply the STRONG-AYA infrastructure for their own projects. Complementary documentation This document is hereafter referred to as the Data Management Plan, or the DMP in short. The DMP is intended for the use of partners of the STRONG-AYA Project (reference number 101057482). The partners are defined in the General Agreement (executed on 18 May 2022, first amendment dated TBC), have executed the Consortium Agreement (20 December 2022), and whose tasks are detailed in the project scope of work in the aforementioned General Agreement. The DMP is to be used complementarily with the STRONG-AYA (governance, legal and operational) ecosystem document – see Deliverable 2.1. STRONG-AYA – No. 101057482 – D3.1 The DMP is to be used complementarily with the STRONG-AYA technical blueprint document – see Deliverable D3.3. The DMP is to be used complementarily with the STRONG-AYA record of processing activity (ROPA) and data protection impact assessment (DPIA) – see Deliverable 2.2. The definitions used in the DMP aligns to terminologies in the forthcoming STRONG-AYA joint data controllership and data processing agreement. 2 Procedure The preliminary COS Data items were developed by consensus among the partners in Work Package 1. A general wish-list of data domains were surveyed among participating clinical institutions and cancer registries. Pragmatically, a more detailed list of data fields were developed as already routinely collected, and would be feasible to collect in future. The COS is intended to grow and scale in future, as partners journey further into this project. With Work Package 2, surveys and unstructured interviews were conducted in order to understand the existing data flows within partner institutions, and potential for scaling up in future. This informs the evaluation of processing activities and data protection impact assessment. This work is ongoing. The principal work of ecosystem description and technical blueprint development was carried out by feasibility surveys and unstructured interviews among partners, and drawing heavily from expertise in successful federated learning projects led by IKNL and University Maastricht. This work is ongoing, however the baseline technical blueprint will be made available online as a an “agile” document and thus incorporates extra functionality as this project progresses. Engagement with a wider network of current and future stakeholders is the work of WP4, and is ongoing. The engagements are essential to understand the needs of the future data contributors and potential data users of the federated ecosystem, thus the technical blueprint and the DMP will evolve hand in hand. This interim DMP was drafted based around the EU Horizon2020 DMP template and the guidance documentation from the international GO-FAIR Initiative. The co-contributors have reviewed this DMP document, and agree on this as an interim DMP, in order proceed further with ecosystem design and technical development, leading towards a finalized DMP at the end of the project as promised. 3 Interim Data Management Plan Core Outcomes Data 3.1.1 Statement of intended purpose of collection The statement of intended purpose has been based on text from the STRONG-AYA (governance, legal and operational) ecosystem report – Deliverable 2.1. Adolescents and young adults (i.e., AYAs) with cancer form a unique group; they face age-specific issues (e.g. infertility, unemployment, financial problems) and decreased quality of life due to cancer and its treatment. Unlike dedicated healthcare services and clinical trials for paediatric cancer patients, AYA-specific healthcare services remain scarce and practice is heterogeneous across Europe. AYAs, who also are at the very core of society and the economy, need access to evidence-based age-appropriate and high-quality dedicated healthcare. STRONG-AYA – No. 101057482 – D3.1 AYA care and research will benefit from collection and pooling of patient-centred data and collaboration among a wide spectrum of stakeholders: patients, healthcare professionals, scientists, and policymakers. The comparisons and descriptive analyses carried out will be used to improve cancer care and treatment at the policy, organisational, service and interpersonal level. 3.1.2 Types of data collected The project intends to make use of wide-ranging patient data, fine details pending the final COS report, in the following general domains. This preliminary list has been provided as a starting point for the consortium to develop and extend and scale up; it is not to be taken as the final and definitive declaration of a finalized COS data set. Determinants and case mix factors - Tumour/disease phenotype (such as staging, site and type of cancer) - Treatment(s) received (such as surgery, chemotherapy, radiotherapy, immunotherapy or combinations thereof) - Biological sex and self-defined gender pronouns - Age at cancer diagnosis - Social determinants (such as socioeconomic grading, ethnicity/cultural background, highest education attained, employment category, domestic status, if there are dependent persons) - Body composition (such as body mass index) - Pre-existing medical conditions - Risk stratification - Use of diagnostic imaging for diagnosis and disease quantification Physiological and clinical outcomes Quality of life outcomes and impacts on activities associated with daily living - Physical functioning, Social functioning, Role functioning, Emotional functioning, Cognitive functioning - Global quality of life - Perceived health status - Perceptions of delivery of care (Satisfaction/patient preference, Acceptability and availability of care, Adherence/ compliance, Withdrawal from treatment, Appropriateness of treatment) - Process, implementation, and service outcomes Health system resource utilization pattern - Economic/Financial hardship - Treatment-related injury/toxicity - Insurance status - Hospitalization and post-treatment health care utilization (e.g. treated as adult or by a paediatric specialist) - Need for further intervention after cancer intervention - Societal burden/carer burden - Medical adverse events and after-effects Mortality - All-cause mortality STRONG-AYA – No. 101057482 – D3.1 3.1.3 Re-use of existing data within this project In STRONG-AYA, primary reliance shall be placed on data obtained during routine cancer care and standard clinical procedures (i.e., real-world data) supplemented by data extractions from previously concluded clinical trials (that might be either observational or experimental in nature). This means STRONG-AYA is principally a DATA SECONDARY USE type of study. Appropriate data pertaining to be abovementioned COS general domains shall be sourced by partner institutions from within their own clinical environment (such as chart audits, electronic health records, quantitative questionnaires), healthcare providers in their care network and from their adjacent national/regional/hospital-based cancer registries. The quality of life domain in the abovementioned data is intended to be sourced via patient-reported information. The COS domains and data fields collected may be extended in scope and scale within the project; the DMP shall be modified accordingly in future versions to capture such changes in scope and scale of data gathering by project partners. One of the intentions in STRONG-AYA is to support the creation, extension and sustainability of institutional workflows to collect high quality COS data well into the foreseeable future. This means that future data collection items need to become implemented into routine clinical care as well as in standard follow-up, and changes also need to be implemented into reporting protocols for cancer registries. 3.1.4 PROMs as a special category of data collection In healthcare, the collection and utilization of Patient Reported Outcome Measures i.e. PROMs is an ever increasingly important source of data about quality of life, impacts on activities related to daily living and perceptions of the utility of care. However, systematic collection of PROMs among the STRONG-AYA partners is presently highly heterogenous, and only some items are being recorded systematically and in all places. The collection and utilization of PROMs shall be an important aspect of STRONG-AYA. Some partners have digital PROMs acquisition systems that contact the patient and solicit electronic responses to PROMs-based quantitative surveys. Some partners do not have any PROMs acquisition systems at all, and other partners rely on paper forms which then have to be collected for digitization. Furthermore, patient engagement will be crucial, as the AYA patient community will end up being the drivers of the design for PROMs collection and PROMs utilization. The STRONG-AYA project is committed to harmonization of electronic PROMs data collection and its integration with as many existing systems as possible, but the exact technical design of such an infrastructure hinges upon the selection of key utilization cases which thus requires extensive patient engagement; this is the principal work of Work Package 4. The PROMs utilization design shall not violate the overall principles laid out in the STRONG-AYA vision, but detailed data management planning pertaining specifically to PROMs remains impossible at the present time. A complete PROMs utilization plan will be included in the definitive version of this DMP to be delivered at end of project. 3.1.5 Data sources Retrospectively collected and real-world data Clinical case mix factors, treatment/intervention information, diagnoses, physiological outcomes, clinical follow-up and mortality will predominantly be sourced from electronic medical health records maintained by the participating institution and/or by extraction from already-closed clinical trials. Social determinants of care, quality of life, impacts on activities of daily living and health system utilization are intended to be extracted from cancer registry databases and supplemented by extraction from alreadyclosed clinical trials. STRONG-AYA – No. 101057482 – D3.1 Prospectively collected data For future, expanded and more definitive versions of COS data fields, the abovementioned institutional electronic health records and cancer registries need to be expanded and updated. This is an operational question that is impossible to answer at the very outset the project, however discussions are ongoing with partner institutions to find out what will be feasible for future data collection. General and specific considerations of the data The Declaration of Helsinki defines “vulnerable groups and individuals“ as those having an increased likelihood of being wronged or of incurring additional harm – in the context of medical research. Generally speaking, vulnerable individuals or populations are those whose decision making capacity and ability to exercise and protect their rights is restricted due to their mental, social, economic or ethnical status, age or health condition or, in other words, those who in one or more ways are at risk of being harmed through the conduct of research. The target population of STRONG-AYA overlaps with subjects that would be commonly deemed as “vulnerable individuals”, specifically persons with incurable diseases, persons who are unemployed or impoverished, persons from ethnic minority groups, and minors. The inclusion of such subjects’ data requires careful considerations in terms of the research activity being carried out as well as protection of identifiable personal characteristics. The in-depth development of governance principles of STRONG AYA will be completed by the end of the first year of the project. This is one of the principal outputs resulting from Work Package 2. An inventory of relevant existing ethical and governance codes and guidelines already present in the various partner institutions will result in the development of the project’s ethical and legal guidelines and framework, and this shall align principles from the different STRONG AYA partners as well as relevant EU governance and ethics principles (such as the ‘European Ethical Principles for Digital Health’ 2022), and with the key principles laid out in the General Data Protection Regulations (GDPR) with its relevant amendments. The following is based on text from the STRONG-AYA (governance, legal and operational) ecosystem report – Deliverable 2.1. The aforementioned principles shall include but are not strictly limited to: - lawfulness, fairness and transparency: data shall be processed lawfully, fairly and in a transparent manner in relation to the data subject; - limitation of purpose: data shall be collected for the specified, explicit and legitimate purposes and not be further processed in a manner that is incompatible with those purposes; further processing for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes shall, in accordance with Article 89(1), not be considered to be incompatible with the initial purposes; - data minimisation: data shall be adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed; - accuracy: data shall be accurate and, where necessary, kept up to date; every reasonable step must be taken to ensure that personal data that are inaccurate, having regard to the purposes for which they are processed, are erased or rectified without delay; - limitation on storage: data shall be kept in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed; personal data may be stored for longer periods insofar as the personal data will be processed solely for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) subject to implementation of the appropriate technical and organisational measures required by this Regulation in order to safeguard the rights and freedoms of the data subject; STRONG-AYA – No. 101057482 – D3.1 • Software tools (e.g., FAIRifcation tools, synthetic data examples, federated infrastructure and federated analysis/learning algorithms) will become an integral part of an established and wellmanaged open-source set of products, so as to maximise re-usability. R1.2. (Meta)data are associated with detailed provenance • All elements associated with a URL or URI are version controlled. • Data provenance is monitored by the data-owning institutions. Insofar as possible, institutional data workflows are scripted and automated and documented in a future partner-specific data workflows report. • Overall responsibility for data lifecycle management (e.g., storage, back-ups, deletions and integrity of underlying data sources) lies with the data providers, and FAIR-specific tooling to update, add and deprecate data will be provided to partners. R1.3. (Meta)data meet domain-relevant community standards • Implementation of RDF and Semantic Web meets domain-specific standards wherever these already exist, particularly for data. • All reasonable steps are taken to be as FAIR as possible whosoever established community standards do not exist. 3.5.5 Data Loss Mitigation Partner institutions that host data, in the federated analysis/federated learning paradigm, keep their own data “on premise” i.e., the COS data items are placed in a secured research and computational environment controlled by the partner institution themselves. There will be no transfer of patient-level data (in any form) from one STRONG-AYA partner to another, for the purpose of the kinds of analyses permitted in the STRONG-AYA framework. The technical blueprint document – see Deliverable 3.3 – requires that on-premise COS data shall be pseudonymized (no native institutional identifiers or equivalents of social security numbers are allowed), minimized (superfluous information that is not needed directly for analysis such as birthdates and postcodes are not allowed) and routinely backed-up as per the institution’s own standard operating procedure pertaining to information technology devices. In event of catastrophic loss of the COS data and its back-up copy, for whatever reason, it must be noted that the original data exists in the primary source systems such as electronic hospital records, clinical trials archives and the cancer registries. Although extremely inconvenient, the COS data extraction procedures put in place by the partner institutions that read from the primary data sources would be automated to the maximum extent possible, therefore it may be re-done from scratch if truly necessary. The likelihood of all systems failing at the same time would be very low, and would suggest a higher priority crisis at the immediate hand than the issue of recovering project data. The COS data will be preserved on-premise by the institutions, for at least the duration of the STRONG-AYA project, as well as beyond the term of the STRONG-AYA project as defined in the legal, governance and business operations framework. 3.5.6 FAIR software STRONG-AYA will take all reasonable steps to make the software in the project to be as OPEN and FAIR as possible. The software resulting from the execution of this project includes : STRONG-AYA – No. 101057482 – D3.1 - Future versions of Vantage6 infrastructure for federated analysis and federated learning - Code for performing federated data analysis and federated epidemiological modelling - Software tools for making structured clinical data FAIR - Interactive dashboards for visualization of summary statistics - Web-based applications for developing and validating epidemiological models in a federated ecosystem Making this software FAIR shall include (but not exclusively limited to) the following : - Programming code will be open source and open access via a widely-used public software repository such as GitHub, along with rich descriptive text informing the intended user how the software shall be installed and operated; - Persistent universal resource locator links (URLs) to the software repositories will always be given in publications and disseminations pertaining to its use in STRONG-AYA results; - The research results that cite the software repository URLs will themselves have persistent Digital Object Identifiers (DOIs) that will be publicly searchable and findable, and the research results (to the maximum extent possible) will be published as open access documents; - The programming code shall principally be in a widely-understood computer language such as Python, Java or C, or a mathematical scripting language such as R, and furthermore each repository shall contain links to software library dependencies (such as a requirements.txt file in Python) to assist reproducibility and re-usability; - Re-use of the software will be widely encouraged using licensing conditions such as Creative Commons By Attribution (with or without Share-Alike option), the GNU Public software license, and other such similar licenses that are widely used by the open-source software community. In addition to the above steps, any epidemiological models developed and validated in the STRONG-AYA project shall be documented (along with its persistent identifiers and URLs) with the AIME (Artificial Intelligence in bioMEdical research) registry for added transparency - https://aime-registry.org/. The models resulting from STRONG-AYA will also be contributed to the newly started FAIRmodels project (https://fairmodels.org/) which shall hold rich metadata on the type of data on which the model has been trained on, how the model was developed and the anticipated limitations of the model’s scope of use. 3.5.7 FAIR services In our case, “services” refer to the operational know-how of creating a consortium around a federated data ecosystem paradigm, which includes: - A formal consortium agreement that defines the general rights, roles and duties pertaining to consortium membership; - A governance and operations framework that defines the work of the consortium, including its process for legal review and ethical review for what the consortium may (and may not) do with the collected data, as well as the humanistic values (e.g. principles of patient empowerment) guiding its conduct; - A technology blueprint defining the hardware and software requirements that constitute the technical parts of the federated infrastructure; - A set of legal mechanisms, such as a joint data controllership agreement and an infrastructure user agreement with third-party service providers, that are necessary to allow federated learning/federated analysis to be performed. - An information model (schema) including a taxonomy/lexicon/terminology defining to the STRONGAYA COS items, but without the data itself. STRONG-AYA – No. 101057482 – D3.1 STRONG-AYA makes the above services open (to the maximum extent possible) and FAIR in the following manner: - Where appropriate, the aforementioned instruments shall have names, addresses and institutions and other potentially identifiable and/or confidential information redacted, such that the resulting document will be made openly accessible (with publicly accessible URL) as a template for future collaborations; - Each of the abovementioned documents will be registered with a persistent DOI, and the entire package of documents comprising the STRONG-AYA services will be listed on a publicly-searchable FAIR objects catalogue (such as Zenodo); - The FAIR catalogue entry shall include rich descriptions about these services, and these descriptions shall remain in the FAIR repository even if the documents themselves are taken offline; - Re-use of the services as templates will be covered by a permissive open access license to be selected by the consortium at the end of the project. 3.5.8 Post-Project Data Sustainability Data in the COS project is made FAIR as a condition of its use during the project, so there are no special costs to make data FAIR after the end of the project. As stated above, one of the intentions in STRONG-AYA is to enable creation, extension and sustainability of institutional workflows to collect high quality COS data well into the foreseeable future and this necessarily implies tools for making the data FAIR on-premises. Since the data stays within the institution that contributes the data, STRONG-AYA methodology must integrate with existing data management procedures as well as data management responsibilities that already exist within partner institutions. As the project, STRONG-AYA is responsibility for the governance, operation and patient rights protection within the federation infrastructure. Costs for long term operation and preservation of the federation infrastructure is part of the governance and business operation plan for the entire consortium. Costs for long-term operation and preservation of the COS data within the participating institutions will be borne directly by that institution itself. Costs for the ongoing delivery of the AI technology will have to be borne by the consortium as a whole, as these technologies permeate further into routine application. However, describing the means by which these organisations might finance or otherwise share costs with the consortium for the maintenance of their COS data will be a valid topic for the governance and business operation plan. 4 Conclusion This deliverable report consists of an interim Data Management Plan that is adequately detailed to act as a temporary roadmap for further development and evolution of the STRONG-AYA project. It will be expected that ongoing consultations between work packages and with stakeholders will provide deeper detail in certain sections of this report, as the project itself takes increasingly concrete shape over the coming years. This Data Management Plan captures the current design and implementation plan for the technology and situation-specific requirements of the project, wheresoever it intersects with data privacy needs of individuals. A finalized Data Management Plan has been promised as a further deliverable near the end of the project. Repository for primary data (annex) Not Applicable.