scieee AI-readable full text Open interactive document viewer

UC 1 Genotypic BigData from FAIR data cohorts in the digital age of plant breeding

Gundala, Ravindra Reddy; Gogna, Abhishek; Chu, Jianting; Fiebig, Anne; Schüler, Danuta; Lange, Matthias; Zhao, Yusheng; Reif, Jochen

Abstract

Genotypic data can be generated using sequencing-by-synthesis (Slatko et al., 2018) (in VCF format) or SNP arrays (e.g. in HapMap format), each potentially producing a distinct FAIR Digital Object (FDO). An FDO is digital data bundled with rich metadata and a permanent distinct ID, making it easier for humans and machines to find, interpret, and reuse. Genotypic FDOs can be created by depositing variant calls, bio-sample information, and raw data in EBI repositories (Figure 1). These genotypic FDOs can then be integrated across data cohorts, where a “data cohort” refers to the collection of diverse data types generated within a single study (Figure 2; Gogna et al., 2025, Accepted). Integration of VCF-based FDOs requires a consistent reference genome (e.g., RefSeqV2.1 for wheat; Zhu et al., 2021). While, HapMapbased FDOs are more difficult to combine due toproprietary probe sequences, uncertain variant positions, and the challenges of variant definition in large or polyploid genomes (Martin et al., 2022). Nevertheless, FDO-driven integration enables building genotypic BigData across cohorts, enabling genomic prediction, GWAS, and QTL mapping, and ultimately accelerating plant breeding.

Full text

Ravindra Reddy Gundala1, Abhishek Gogna1, Jianting Chu1, Anne Fiebig1, Danuta Schueler1, Matthias Lange1, Yusheng Zhao1and Jochen C. Reif1,* 1. Leibniz Institute of Plant Genetics and Crop Plant Research (IPK), 06466 Seeland, Germany *. Corresponding author: Jochen C. Reif ([email protected]) UC1–Genotypic BigData from FAIR data cohorts in the digital age of plant breeding Summary: Genotypic data can be generated using sequencing-by-synthesis (Slatko et al., 2018) (in VCF format) or SNP arrays (e.g. in HapMap format), each potentially producing a distinct FAIR Digital Object (FDO). An FDO is digital data bundled with rich metadata and a permanent distinct ID, making it easier for humans and machines to find, interpret, and reuse. Genotypic FDOs can be created by depositing variant calls, bio-sample information, and raw data in EBI repositories (Figure 1). These genotypic FDOs can then be integrated across data cohorts, where a “data cohort” refers to the collection of diverse data types generated within a single study (Figure 2; Gogna et al., 2025, Accepted). Integration of VCF-based FDOs requires a consistent reference genome (e.g., RefSeqV2.1 for wheat; Zhu et al., 2021). While, HapMapbased FDOs are more difficult to combine due to proprietary probe sequences, uncertain variant positions, and the challenges of variant definition in large or polyploid genomes (Martin et al., 2022). Nevertheless, FDO-driven integration enables building genotypic BigData across cohorts, enabling genomic prediction, GWAS, and QTL mapping, and ultimately accelerating plant breeding. Figure 2: An example workflow for genotypic data integration. Recoded and curated HapMap-based SNP array FDOs are integrated according to the platform and oligo source. HapMap datasets are first converted into VCF format to correct platformspecific biases, enabling their seamless integration with VCF-based datasets across cohorts. This process generates large-scale genotypic data suitable for downstream genomic analyses. Figure 1: Whole genome re-sequence (WGS) data FAIR Digital Object (FDO) creation pipeline at IPK for the use case GENEBANK3.0. References: Gogna et al., 2025. Accepted Martin et al., 2022. https://doi.org/10.1093/nar/gkac958 Slatko et al., 2018. https://doi.org/10.1002/cpmb.59 Zhu et al., 2021. https://doi.org/10.1111/tpj.15289 (Adapted from: The AGENT consortium et al. AGENT Guidelines for dataflow. (2024) https://doi.org/10.5281/zenodo.12625359)