scieee AI-readable full text Open interactive document viewer

FAIR Research Data Management - Getting started with putting FAIR RDM into practice

Thorpe, Deborah Ellen; Brinkman, Loek; van den Berk, Michelle; van Horik, René

Abstract

These training materials are part of a five-part course aimed at early career researchers on FAIR Research Data Management (RDM). We share these materials so that research data professionals can reuse them in their instructions or training sessions. The main theme of this course are the FAIR principles: sharing research data that are Findable, Accessible, Interoperable and Reusable.The FAIR RDM course is intended to be flexible and modular: we invite users to (re)use and adapt those parts that are suitable for their audience.Sessions 1-3 of this course are aimed at absolute beginners: that is, those who likely have knowledge of research processes in general but are new to research data management and the FAIR principles. Sessions 4-5 are aimed at an intermediate level: those who have completed the first three modules and/or already have some existing knowledge of FAIR RDM. The training session ‘Getting started with putting FAIR RDM into practice’ was piloted in September 2024 as the third session of the 'FAIR Research Data Management' 5-session course (beginner and intermediate level). The course was developed and piloted as part of the PATTERN project (https://www.pattern-openresearch.eu/). In ‘Getting started with putting FAIR RDM into practice’ learners are introduced to the different types of metadata and documentation and why they are important in relation to the FAIR Principles. In the lecture that follows, there is an overview of the common file formats for different types of data, which file formats support FAIR and the difference between open and proprietary file formats. Finally, there is a beginner-level introduction to the basics of data sharing and repositories, demonstrating the role of repositories in FAIR. Learners are introduced to Re3data and FAIRsharing as tools for finding the appropriate repository for their data. Following the lecture/discussion component of the session, there is an exercise on reading and evaluating README files, as a simple but effective way to create documentation for data. Finally, there is the project work for this section, where learners continue to apply their knowledge to a project that unifies several sub-areas of FAIR research data management. Though learners will take away knowledge that is broadly applicable across the disciplines, there is a focus on discipline specific learning through the presentation of READMEs in different domain areas; through the exercises; and by offering discipline-specific choices for project work. Learning Outcomes: By the end of this session, learners will be able to 1. Describe types of documentation and metadata2. Outline why metadata is important in relation to the FAIR Principles3. Give an example of a type of data documentation and briefly describe the type of information it might contain4. Suggest which file formats support FAIR data5. Outline what the differences are between open and proprietary formats6. Describe what a repository is and begin to explain a repository’s role in FAIR Project work: An important component of these training sessions was project work that was conducted on the Projects platform (https://pattern.projects.directory/), which you will see references to throughout the slideshows in this series. We asked learners to do some exercises that are based on real research projects that produced and archived data some time in the past. Eight 'use cases' were created, and participants chose one that mached their interests to work on for the duration of the course. From the first session onwards, they then worked on these 'use cases' in the Projects platform, with regular 'check ins' with other learners during the live training sessions. The eight use cases have been uploaded to Zenodo separately. File overview:20240910_PATTERN_FAIR_RDM_Session3_Slides: The central slideshow for this 2 hour training session20240903_FAIR_RDM_Session3_Exercise.docx: Exercise that was completed in groups during the training session20240903_PATTERN_FAIR_RDM_Session3_SessionPlan: the document that we created to plan and manage the training session, which we believe will be useful for potential reusers of the content Slideshows are uploaded in .pptx and .pdf format and text documents are uploaded in both .docx and .pdf Related records: Session 1: 10.5281/zenodo.15310232 Session 2: 10.5281/zenodo.15310356 Session 4: 10.5281/zenodo.15310506 Session 5: 10.5281/zenodo.15310556 FAIR RDM Use Cases: 10.5281/zenodo.15316306

Full text

Exercise: Read a README! Introduction and recap on README Files: ● A README file is a simple way to create documentation for a file(s) that is valuable for reuse ● A README should be in an open .txt format that can be opened by anyone in a simple file viewer ● The purpose of a README file is to provide information about a file that can help to ensure that it can be correctly interpreted by yourself at a later date, or by others when sharing or publishing data ● There are other forms of documentation, e.g. Lab notebooks, codebooks, a README is just one type and is the most straightforward to produce See: ‘Guide to writing “readme” style metadata’, Cornell Data Services, https://data.research.cornell.edu/data-management/sharing/readme An exercise on evaluating README Files: ▪ As an example of ‘best practice’, take a look at the README Template below. This template has been extracted from Cornell Data Services, ‘Guide to writing “readme” style metadata’, https://data.research.cornell.edu/data-management/sharing/readme/ (take 5 minutes to examine the template) ▪ Choose one of the datasets in the table below, - Skim read through the information about the dataset (metadata) that you see in the Repository - find the README file among the data files, download it, and read it (take 15 minutes) Datasets to choose from (choose one) Albertson, Lindsey (2022). Influence of beaver mimicry restoration on habitat availability for fishes, including Arctic grayling (Thymallus arcticus) [Dataset]. Dryad. https://doi.org/10.5061/dryad.47d7wm3fq Rachel F Marek et al, Dataset for airborne PCBs and OH-PCBs inside and outside urban and rural U.S. schools, Iowa Research Online, https://doi.org/10.25820/data.002114 Ferguson, Jake M; Fieberg, John R; McCartney, Michael A.; Blinick, Naomi S.; Schroeder, Leslie. (2019). Data and R code to support: Estimating densities of zebra mussels (Dreissena polymorpha) in early invasions using distance sampling. Retrieved from the Data Repository for the University of Minnesota (DRUM), https://doi.org/10.13020/d6hc-bw36. C.T. Elmelund, 2020, "Transcriptions of the subscriptions to 2 Timothy in 485 Greek manuscripts", https://doi.org/10.17026/dans-xk3-dabp, DANS Data Station Social Sciences and Humanities, V2 Exercise: 'Getting started with putting FAIR RDM into practice': Session 3 of PATTERN training in FAIR RDM. DOI: 10.5281/zenodo.15310456 Milan van Lange; Carlijn Keijzer; Annelies van Nispen, 2023, "First-Hand Accounts of War: War Letters (1935-1950) from NIOD Digitised", https://doi.org/10.17026/SS/UUVUW2, DANS Data Station Social Sciences and Humanities, V3, UNF:6:MaYfqUTK0jbjYFIjYLHr0Q== [fileUNF] P. Heijnen, 2016, "Retail gasoline prices in the Netherlands 2005 - 2011", https://doi.org/10.17026/dans-25c-56vs, DANS Data Station Social Sciences and Humanities, V3, UNF:6:J5w+NIDIu41uXEOZ8q8OVA== [fileUNF] ▪ Discussion (5-10 minutes) 1) What information can you find in the README file that would help with understanding and possibly reusing the data? 2) What could be improved/what information are you missing? Take home message: We should always be documenting our research as we go – it is difficult to go back and capture details later! Create README files for your data file(s) as early as possible Any others? Exercise: 'Getting started with putting FAIR RDM into practice': Session 3 of PATTERN training in FAIR RDM. DOI: 10.5281/zenodo.15310456