BIDS-flux Distribits 2025 presentation
Full text
BIDSflux building a data operations platform for open-science Basile Pinsard, PhD & the CPIP IT team & the NeuroMod team & CRIUGM Open-Science team
Context: deep neuroimaging at scale CCN 2025
Context: NeuroMod setup
Context: instruments to analysis-ready data
FAIR data management of file-based structures Everything is stored in modular datalad repos: - track provenance of sourcedata with submodules (Yoda) -datalad (container-)run to track processes generating/modifying data https://handbook.datalad.org/en/latest/basics/101-127-yoda.html
scaling?! Collect data… download it 3y later, realize missing info, design flaws and acquisition issues. n different students, run x times the pipelines y/z/ on different versions Data-sharing: put back the pieces together for release. Curate, prepare, release and deploy data continuously and automatically FAIR principles (Findability, Accessibility, Interoperability, and Reusability) Provide researcher simple access to most recent analysis-ready data.
Ethics Ethics Plan Plan Ethics Collect Prepare Share Preserve Collect Prepare Share Preserve Open-Data/Open-Science lifecycle Design Funding - Log - Extract from instruments - Store - Backup - Aggregate - Standardize - Anonymize - Comply protocol - QC/review - Processing - Governance - Sovereignty Ethics Plan - Specs - SOPs - Piloting Collect Prepare Share - FAIR - Security - Access mngt. Preserve - Archive Experimenter Registered report Plan Analyze Report Gather Data - Discover - Apply - Access - Track version - Reproducible env. - Track processes - Version outputs - reproducible notebooks / preprints User Publish - Data mngt. plan
A Maturity Model for Operations in Neuroscience Research (Johnson et al. 2023, preprint) CNeuroMod is here
DataOps feedback loop to data collection Iteratively add more tests/feedback when errors occur: - Add instruments checks: eg. export uncombined data on localizer to check for receive coil problems. - Add new feedback loop: artifacts detected -> @mri-operators please export raw k-space data
DataOps as-a-service
Testing of data: example of protocol compliance (WIP) - Session structure: sequences (optional), order, timing - Check BIDS metadata against a predefined JSON-schema - Multi-scanner/site design: allow sharing sub-schema by Manufacturer/Model/Scanner https://github.com/UNFmontreal/forbids
Open-science is hard: data user is left on its own - messy part
Open-science is hard: lower barriers End-to-end framework
Open-science is hard: reduce friction along the full lifecycle
File-based standard to other formats.
Not quite yet! This is local!
What about multi-site studies? Distributed DataOps Heterogenous Canadian human-data sharing landscape: - Federal laws - Provincial laws - First-Nations info governance. - Institutions … Flexible sharing mechanisms. Local sovereignty/storage/control. Federated metadata indexes. Sharing of less sensitive aggregates. Federated analysis.
Summary Datalad empowers: ●Scalable distributed DataOps on file-based data standards. ●Flexible Open-Science, even under sovereignty and governance constraints, through discovery of federated rich metadata ●End-to-end FAIR data workflow and analysis that enable meta-science. What we want to develop: ●Generic glue for pipelines CI/CD orchestration on modularized datasets. ●Distributed and federated infrastructure. ●Standalone generic tools that operate on data standards and Datalad spec. ●Infrastructure and processes to reduce barriers and friction to collaborative Open-Science. 🚀