scieee AI-readable full text Open interactive document viewer

Data derived from sex-specific recombination rate analysis in cattle

Wittenburg, Dörte

Abstract

Frequencies of sex-specific recombination events have been studied in eight European cattle breeds comprising dairy, dual-purpose and beef breeds (German Holstein, Swiss Holstein, Brown Swiss, German/Austrian Fleckvieh, Original Braunvieh, Simmental, Limousin, Angus). The data comprise physical and genetic-map coordinates of about 40K single nucleotide polymorphisms as well as various summary statistics.

Full text

Description of data derived from sex-specific recombination rate analysis in cattle N. Melzer and D. Wittenburg Research Institute for Farm Animal Biology (FBN) Date: 18.12.2025 Version: 2.0 Contact: [email protected] 1. Data sets generated applying the hsrecombi pipeline Data sets were generated for the Clarity App (v3.0.0, GitHub) using the hsrecombi pipeline (v2.0.1, GitHub) with the following approaches included: deterministic approach (hsphase; v2.0.2, CRAN), likelihood-based approach (hsrecombi; v1.0.1, CRAN) and Hidden-Markov-Model (HMM) based approach specific for male, female and average (LINKPHASE3; v2.0.0, GitHub). Breed-specific data were stored in breed-specific folders. Exception: The HMM-based approach is not included for the Holstein-DE data. 1.1 geneticMap.Rdata The Rdata file contains a data frame termed “geneticMap” with relevant genetic map information of all markers and chromosomes derived from the different approaches. Header description: “geneticMap_header.csv” 1.2 genetic_map_summary.Rdata The Rdata file contains a data frame termed “genetic_map_summary” with descriptive statistics and key metrics for every chromosome and overall chromosomes derived from the different approaches. Header description: “genetic_map_summary_header.csv” 1.3 adjacentRecRate.Rdata The Rdata file contains a data frame termed “adjacentRecRate” with information about recombination rate between adjacent markers on all chromosomes derived from the HiddenMarkov-Model approaches (male, female, average) and the deterministic approach. Header description: “adjacentRecRate_header.csv” 1.4 bestmapfun.Rdata The Rdata contains a matrix termed “out” with information about estimated parameters of genetic-map functions (Haldane scaled, Rao, Felsenstein, and Liberman & Karlin) and mean squared errors of curves fitted to recombination rates that were derived from the likelihoodbased approach for all chromosomes. Header description: “bestmapfun_header.csv” 1.5 curve-short-<chr>.Rdata The file contains relevant information of four used genetic-map functions (Haldane scaled, Rao, Felsenstein, and Liberman & Karlin) fitted to recombination rates that were derived from the likelihood-based approach for each chromosome (for more details see Melzer et al. 2023). Each file contains a list termed “store” with four elements: 1. Matrix (numeric values): genetic distance in Morgan (dist_M) and recombination rate (theta) between markers. Only a reduced number of data points are provided; a data point is a pair of recombination rate and genetic coordinate ordered according to a vectorized triangular matrix of SNP identifiers. 2. Matrix (numeric values): x-values of applied genetic-map functions (order: Haldane scaled, Rao, Felsenstein, and Liberman & Karlin). 3. Matrix (numeric values): y-values of the applied genetic-map functions (order: Haldane scaled, Rao, Felsenstein, and Liberman & Karlin). 4. Numeric value: percentage of marker pairs contained in matrix. Note: no specific header files were created. 2. Additional data sets Additional data sets contain information about markers that are putatively misplaced in the underlying reference genome ARS-UCD1.2 as well as information about chromosome regions that were difficult to assemble in general. 2.1 generalProblematicRegions.Rdata The file contains a tibble data frame termed “generalProblematicRegions” with information about problematic chromosome regions that are recommended to be removed from genome-based analysis regardless of the SNP panel used (Qanbari et al. 2022). Header description: “generalProblematicRegions_header.csv” 2.2 misplacedMarkers.Rdata The file contains a tibble data frame termed “misplacedMarkers with information about 65 SNP markers with strong evidence for being misplaced in the bovine reference genome based on the analysis of male recombination rate in German Holstein cattle (Qanbari et al. 2022). Header description: “misplacedMarkers_header.csv” 2.3 OverviewBreeds.Rdata The file contains a data frame termed “OverviewBreed” with information about number of genotyped individuals and number of paternal half-sib families for all breeds included. Header description: “OverviewBreed_header.csv” 2.4 misplaced_all_breeds.Rdata The file contains a tibble data frame termed “tab” with candidate misplaced markers detected for the breeds: Brown Swiss, Braunvieh, Holstein (CH), Simmental, Angus, Limousin and Fleckvieh. Header description: “misplaced_all_breeds_header.csv”. Acknowledgments We gratefully acknowledge the support of the: • Association for Bioeconomy Research (FBF, Bonn, Germany) as representative of German Holstein cattle breeders for participating in this project and the German Evaluation Center (VIT, Verden, Germany) for composing the Holstein data. We thank ZuchtData (Vienna, Austria) for providing the German/Austrian Fleckvieh data. • Swiss cattle breeders, represented by Mutterkuh Schweiz, Braunvieh Schweiz and Genossenschaft Swissherdbook Zollikofen, for participating in this project. We thank Qualitas AG (Zug, Switzerland) for preparing and providing the genotype data.