scieee AI-readable full text Open interactive document viewer

We can only have good science if we have good data - A PSDI webinar

Partridge, Matthew

Abstract

The Physical Sciences Data Infrastructure (PSDI) aims to accelerate research in the physical sciences by providing a data infrastructure that brings together and builds upon the various data systems researchers currently use. Our webinar series provide updates on our exploratory pathfinder work, and on relevant tools and technologies that have been developed by members of our community. This record contains a more detailed version of the presentation that was presented at the "We can only have good science if we have good data" webinar presented by Dr Matthew Partridge on 16th October 2025. https://www.psdi.ac.uk/event/webinar-data-auditing/ Abstract: This webinar explores the data auditing process behind the Physical Chemistry Properties data collection. It highlights surprising errors found in trusted sources, the methodologies that were explored, and includes a very persuasive argument for why auditing isn’t just housekeeping – it’s key for building reliable data collections without which we won’t have reliable science.

Full text

https://www.psdi.ac.uk/ We can only have good science if we have good data 16th October 2025 Dr Matthew Partridge UK funded project through the UKRI Digital Research Infrastructure theme (DRI) via EPSRC Building a UK, Physical Science, Data Infrastructure Focused on delivering a number of key services including Platform & Services: such as cross-source data search, and format conversion Data sources: PSDI provides access to over 20 different databases and repositories of physical sciences data Tools: Downloadable applications (e.g. MagresView, data conversion library) that can run on local systems to process and transform data Guidance & Training: Resources including knowledge base articles, FAQs, tutorials and Moodle-based course What is PSDI? •Develop methods for creating and managing datasets. •Desing flexible database structures. •Apply shared metadata standards. Building Robust Data Foundations •Combine data from repositories and publications. •Align naming, units, and identifiers. •Apply error and duplication checks. Aggregating and Standardising Data Sources •Curate data for research and modelling. •Include provenance and quality scores. •Ensure open, FAIR-compliant access. Delivering ApplicationReady Data Collections •Showcase real research applications. •Assess performance and scalability. •Refine methods through feedback. Demonstrating through Case Studies Building Data Collections Physical Chemistry Properties Data Collection (PChProp) Physical Chemistry Properties Data Collection (PChProp) Cherqaoui1994 Kim2024 BioQuest WikiData Bergstrom2003 Williams2015 BigSolDB SolProp Boobier2020 Delaney2004 Llompart2024 Lowe2023 Meng2022 DDB2023 IUPAC Sander2023 Physical Chemistry Properties Data Collection (PChProp) Cherqaoui1994 Kim2024 BioQuest WikiData Bergstrom2003 Williams2015 BigSolDB SolProp Boobier2020 Delaney2004 Llompart2024 Lowe2023 Meng2022 DDB2023 IUPAC Sander2023 Aim is to make the collection reliable, traceable, usable, and consistent We call this data auditing To do this we need to answer the following questions What is the source of this data? Can we use this data? Is the data complete and understood? Checking for ‘good’ data Peer review evaluates the scientific merit and originality of research, Auditing checks compliance with predefined standards or procedures. Peer reviewers are subject-matter experts assessing content Auditors may not be domain specialists but are focused on quality assurance or regulatory frameworks. Peer review aims to improve scholarly output Auditing aims to maintain the quality of scholarly output Auditing is not peer review BigSolDB SolProp Boobier2020 Delaney2004 Llompart2024 Lowe2023 Meng2022 What is the source of this data? We do source overlap checking via categorisation We compare incoming data with the collection for overlap Highly probable •Matching molecular IDs • Matching sources/references •Identical values Probable •Matching molecular IDs •Matching values • Unknown source/reference Possible •Matching molecular IDs •Closely matching values (within rounding) • Unknown source/references What is the source of this data? What is the source of this data? Can we use this data? Being online doesn’t automatically give you free use Academic fair use and republication are very different things Older (pre-2000) work rarely has licencing information Creative Commons zero (CC0) is the best to look for When in doubt, get consent from originating author Always provide citations Can we use this data? Is the data complete & understood? Details matter Every subject has specific details Not always possible to capture everything but planning is critical Key considerations Experimental conditions Outliers / Impossible values Accidental replicates Accidental replicates Missing values Is the data complete & understood? Folder Data set metadata Describes data about data files We want a fairly basic metadata minimum Author/source (already covered) Date Format Helps with ingest and tracking We never remove metadata but translate this into the collection to ensure provenance is maintained Describes data about individual data We have required a desirable Required Molecular ID (e.g. SMILES) Property value (e.g. 12) Unit values (e.g. Celcius) Desirable More molecular descriptors Experimental conditions Again, we never remove metadata but translate this into the collection to ensure provenance is maintain ID InChI Solubility SMILEScurated SD Group Dataset Composition Error Charge SMILES LogS Compound SMILES InChi Temp Solubility Source LogS Data metadata Is the data complete & understood? Data Submission Provenance Auditing Rights Assessment Metadata Checking Dataset Classification Duplicate & Outlier Reconciliation Identifier Standardisation Final Inclusion Checking for ‘good’ data The full process Auditing the audit Like the data the audit process must be traceable Every step has a documentary trail, and we use unique IDs to track that Any data not included is logged Any data included and changed/fixed is logged Any data sets not included are logged If the audit process changes, we can retrospectively look back and reassess both data in the collection and data not included Auditing the audit Always ask What is the source of this data? Can we use this data? Is the data complete and understood? Not every data set has a simple answers and there is some judgement needed If in doubt, keep detailed records and notes! Just like the source data the choices behind a collection should be verifiable Webinar Summary Thank-you, and questions [email protected]