UC 1 Genotypic BigData from FAIR data cohorts in the digital age of plant breeding
Abstract
Genotypic data can be generated using sequencing-by-synthesis (Slatko et al., 2018) (in VCF format) or SNP arrays (e.g. in HapMap format), each potentially producing a distinct FAIR Digital Object (FDO). An FDO is digital data bundled with rich metadata and a permanent distinct ID, making it easier for humans and machines to find, interpret, and reuse. Genotypic FDOs can be created by depositing variant calls, bio-sample information, and raw data in EBI repositories (Figure 1). These genotypic FDOs can then be integrated across data cohorts, where a “data cohort” refers to the collection of diverse data types generated within a single study (Figure 2; Gogna et al., 2025, Accepted). Integration of VCF-based FDOs requires a consistent reference genome (e.g., RefSeqV2.1 for wheat; Zhu et al., 2021). While, HapMapbased FDOs are more difficult to combine due toproprietary probe sequences, uncertain variant positions, and the challenges of variant definition in large or polyploid genomes (Martin et al., 2022). Nevertheless, FDO-driven integration enables building genotypic BigData across cohorts, enabling genomic prediction, GWAS, and QTL mapping, and ultimately accelerating plant breeding.