scieee AI-readable full text Open interactive document viewer

From Data to Credits: Using ReadMe, Markdown, and Dublin Core for Better Documentation

Dockhorn, Ron

Abstract

Effective documentation is crucial for good research data management (RDM) enabling reproducibility, data sharing, and collaboration. A common standard for information sharing in data publication is to write text files as generic “ReadMe”s. Using the markup language Markdown and the Dublin Core vocabulary allows for formatted and structured documentation. The enrichment of datasets with these metadata, including licenses such as Creative Commons, ensures findability, attribution, and reusability in open science publication, as well as easy-to-understand descriptions for research partners. Furthermore, utilizing parsers to extract and validate metadata can streamline the documentation process. I will provide some practical tips and examples for implementing these concepts establishing a robust data life cycle in your research and facilitating data sharing and collaboration. Link to README.md: https://doi.org/10.5281/zenodo.14848834Link to Markdown-to-JSON-Parser: https://doi.org/10.5281/zenodo.14942696 and https://github.com/Bondoki/ParsingMetadataMD2JSON

Full text

CRC 1415 - Chemistry of Synthetic Two-Dimensional Materials From Data to Credits: Using ReadMe, Markdown, and Dublin Core for Better Documentation README.md Structure of READMEs Motivation Challenge: •Researchers lack standardized documentation practices in RDM, impeding data sharing, collaboration and open science. Insufficient structured documentation and appropriate metadata hinders the findability, attribution, and reusability of datasets. • Goal: •Facilitate common standard for data description using generic "ReadMe" text files for human and machine readability. Utilizing Markdown and Dublin Core vocabulary for formatted and structured documentation for interoperability.[1] Implementing parsers to extract and validate metadata[2] streamlines the documentation process for robust data life cycle. References Dr. Ron Dockhorn¹ 0000-0002-5268-5430 ¹ CIDS - Center for Interdisciplinary Digital Sciences Informationsdienste und Hochleistungsrechnen (ZIH) Technische Universität Dresden • • Basics: •A "ReadMe" file gives a brief description of your project, helping you and others understand what's in your folder or files. The main folder should have a general ReadMe file, and each subfolder and dataset should have its own descriptive ReadMe. Use a simple ASCII text file to describe your work, avoiding formats like .docx or .odt by using the easy-to-read Markdown language. Name file ReadMe.md including appropriate file extension. • • • Contents[1,2]: •Include important details about how data is collected, processed, and analyzed during the research process. At a minimum, the ReadMe should state: Title, Creator, Contact information, File naming convention, Path/URL, License, Description. 5W1H+R (What? When? Where? Who? Why? How? + Relationship) give guidance with additional focus on data linkage.[3] Markdown headers with Dublin Core keywords allow for organized sections, consistent style, and structured documentation. • • • Advantages: •Minimal Standards: Ensure consistency with known keywords. Structured Layout: Clear organization including relational identifiers similar to RDF-schema enhancing linked open data. Interoperability: Simple parsing and machine (LLM)-processing between sections facilitate interoperability across different systems. Compatibility: Utilizes a Zenodo-based[4] JSON file for easy further processing e.g. create database catalog or enrich dataset entries. • • • MD [1] R. Dockhorn, "From Data to Credits: Using ReadMe, Markdown, and Dublin Core for Better Documentation", Zenodo (2025) https://doi.org/10.5281/zenodo.14848834 [2] R. Dockhorn, ParsingMetadataMD2JSON (v1.0.0), Zenodo (2025) https://doi.org/10.5281/zenodo.14942696 https://github.com/Bondoki/ParsingMetadataMD2JSON [3] P. Subramaniam, Y. Ma, C. Li, I. Mohanty, R.C. Fernandez, "Comprehensive and Comprehensible Data Catalogs: The What, Who, Where, When, Why, and How of Metadata Management", ArXiv (2021) https://doi.org/10.48550/arXiv.2103.07532 [4] https://developers.zenodo.org/#representation https://developers.zenodo.org/#github ReadMe CONTENT