https://www.psdi.ac.uk/ We can only have good science if we have good data 16th October 2025 Dr Matthew Partridge
UK funded project through the UKRI Digital Research Infrastructure theme (DRI) via EPSRC Building a UK, Physical Science, Data Infrastructure Focused on delivering a number of key services including Platform & Services: such as cross-source data search, and format conversion Data sources: PSDI provides access to over 20 different databases and repositories of physical sciences data Tools: Downloadable applications (e.g. MagresView, data conversion library) that can run on local systems to process and transform data Guidance & Training: Resources including knowledge base articles, FAQs, tutorials and Moodle-based course What is PSDI?
•Develop methods for creating and managing datasets. •Desing flexible database structures. •Apply shared metadata standards. Building Robust Data Foundations •Combine data from repositories and publications. •Align naming, units, and identifiers. •Apply error and duplication checks. Aggregating and Standardising Data Sources •Curate data for research and modelling. •Include provenance and quality scores. •Ensure open, FAIR-compliant access. Delivering ApplicationReady Data Collections •Showcase real research applications. •Assess performance and scalability. •Refine methods through feedback. Demonstrating through Case Studies Building Data Collections
Physical Chemistry Properties Data Collection (PChProp)
Physical Chemistry Properties Data Collection (PChProp) Cherqaoui1994 Kim2024 BioQuest WikiData Bergstrom2003 Williams2015 BigSolDB SolProp Boobier2020 Delaney2004 Llompart2024 Lowe2023 Meng2022 DDB2023 IUPAC Sander2023
Physical Chemistry Properties Data Collection (PChProp) Cherqaoui1994 Kim2024 BioQuest WikiData Bergstrom2003 Williams2015 BigSolDB SolProp Boobier2020 Delaney2004 Llompart2024 Lowe2023 Meng2022 DDB2023 IUPAC Sander2023
Aim is to make the collection reliable, traceable, usable, and consistent We call this data auditing To do this we need to answer the following questions What is the source of this data? Can we use this data? Is the data complete and understood? Checking for ‘good’ data
Peer review evaluates the scientific merit and originality of research, Auditing checks compliance with predefined standards or procedures. Peer reviewers are subject-matter experts assessing content Auditors may not be domain specialists but are focused on quality assurance or regulatory frameworks. Peer review aims to improve scholarly output Auditing aims to maintain the quality of scholarly output Auditing is not peer review
BigSolDB SolProp Boobier2020 Delaney2004 Llompart2024 Lowe2023 Meng2022 What is the source of this data?
We do source overlap checking via categorisation We compare incoming data with the collection for overlap Highly probable •Matching molecular IDs • Matching sources/references •Identical values Probable •Matching molecular IDs •Matching values • Unknown source/reference Possible •Matching molecular IDs •Closely matching values (within rounding) • Unknown source/references What is the source of this data?
What is the source of this data?
Can we use this data?
Being online doesn’t automatically give you free use Academic fair use and republication are very different things Older (pre-2000) work rarely has licencing information Creative Commons zero (CC0) is the best to look for When in doubt, get consent from originating author Always provide citations Can we use this data?
Is the data complete & understood?
Details matter Every subject has specific details Not always possible to capture everything but planning is critical Key considerations Experimental conditions Outliers / Impossible values Accidental replicates Accidental replicates Missing values Is the data complete & understood?
Folder Data set metadata Describes data about data files We want a fairly basic metadata minimum Author/source (already covered) Date Format Helps with ingest and tracking We never remove metadata but translate this into the collection to ensure provenance is maintained
Describes data about individual data We have required a desirable Required Molecular ID (e.g. SMILES) Property value (e.g. 12) Unit values (e.g. Celcius) Desirable More molecular descriptors Experimental conditions Again, we never remove metadata but translate this into the collection to ensure provenance is maintain ID InChI Solubility SMILEScurated SD Group Dataset Composition Error Charge SMILES LogS Compound SMILES InChi Temp Solubility Source LogS Data metadata
Is the data complete & understood?
Data Submission Provenance Auditing Rights Assessment Metadata Checking Dataset Classification Duplicate & Outlier Reconciliation Identifier Standardisation Final Inclusion Checking for ‘good’ data The full process
Auditing the audit
Like the data the audit process must be traceable Every step has a documentary trail, and we use unique IDs to track that Any data not included is logged Any data included and changed/fixed is logged Any data sets not included are logged If the audit process changes, we can retrospectively look back and reassess both data in the collection and data not included Auditing the audit
Always ask What is the source of this data? Can we use this data? Is the data complete and understood? Not every data set has a simple answers and there is some judgement needed If in doubt, keep detailed records and notes! Just like the source data the choices behind a collection should be verifiable Webinar Summary
Thank-you, and questions
[email protected]