scieee AI-readable full text Open interactive document viewer

Genetic diversity, population structure and linkage disequilibrium in Nordic spring barley (Hordeum vulgare L. subsp. vulgare)

Bengtsson, Therese,The PPP Barley Consortium,Manninen, Outi,Jahoor, Ahmed,Orabi, Jihad

Full text

RESEARCH ARTICLE Genetic diversity, population structure and linkage disequilibrium in Nordic spring barley (Hordeum vulgare L. subsp. vulgare) There ´se Bengtsson .The PPP Barley Consortium .Outi Manninen . Ahmed Jahoor .Jihad Orabi Received: 29 September 2016 / Accepted: 24 January 2017 ÓThe Author(s) 2017. This article is published with open access at Springerlink.com Abstract Genetic diversity, population structure and genome-wide linkage disequilibrium (LD) was estimated in Nordic spring barley (Hordeum vulgare L. subsp. vulgare) by genotyping 180 breeding lines with 48 SSR markers and 7842 high-confidence SNPs using the Illumina Infinium 9K assay. In total 6208 SNPs were polymorphic and selected for further statistical analysis. A Mantel test revealed a strong positive correlation with a Pearson’s correlation coefficient (r) of 0.86, between the estimates of genetic distances based on SSR and SNP data. Population structure analysis identified two groups with a clear ancestry and one group with an admixed ancestry. The groups were primarily separated based on row-type and geographical origin. Average LD for the whole population decayed below a critical level of r 2 =0.20 within a range of 0–4 cM. To avoid confounding effects of the strong population structure, LD decay for the different groups was analysed separately and ranged from 0 to 12 cM. A slower LD decay was found within the two-rowed lines compared to the six-rowed lines and the two-rowed lines originating from the northern part, which could be the result of strong selection for malting quality and yield in the southern part. No large difference in genetic diversity was observed between population sub-groups, but differences at certain chromosomal regions were evident. Keywords Hordeum vulgare L. Linkage disequilibrium Microsatellites Molecular markers  Plant breeding SNP array Introduction Cultivated barley (Hordeum vulgare L. subsp. vulgare) is one of the most important crops in the Nordic countries covering a total area of 1.56 million ha in 2014 (http://faostat3.fao.org). Nordic barley breeding began more than 100 years ago, starting by selections in the landrace gene pool to the present modern elite cultivars developed by crosses between pure lines of advanced material (Kolodinska Brantestam et al. 2004; Fischbeck 1992). Exotic sources are used in modern breeding, but then primarily as donors of single resistance genes against major diseases Electronic supplementary material The online version of this article (doi:10.1007/s10722-017-0493-5) contains supplementary material, which is available to authorized users. T. Bengtsson (&)A. Jahoor Department of Plant Breeding, Swedish University of Agricultural Sciences, Box 101, 230 53 Alnarp, Sweden e-mail: [email protected] O. Manninen Boreal Plant Breeding Ltd, Myllytie 10, 31600 Jokioinen, Finland A. Jahoor J. Orabi Nordic Seed A/S, Kornmarken 1, 8464 Galten, Denmark 123 Genet Resour Crop Evol DOI 10.1007/s10722-017-0493-5 (Melchinger et al. 1994; Weibull et al. 2003). Today’s cultivation of genetically uniform cultivars is raising concerns about loss of genetic diversity. A previous study of the genetic diversity in spring barley germplasm in the Nordic and Baltic region reported a significant decrease of genetic diversity in the spring barley from southern parts of the investigated region in the middle of the twentieth century, but not in the spring barley from the northern parts (Kolodinska Brantestam et al. 2007). Likewise, no signs of genetic erosion were observed in a recent study of genetic diversity for barleys from the northern European area over a hundred years of barley breeding (Rajala et al. 2016). This highlights the importance of knowledge regarding the level of genetic diversity in breeding material, since it enables the detection of any changes in diversity that might lead to genetic erosion. Several types of molecular markers such as amplified fragment length polymorphism (AFLP), random amplified polymorphic DNA (RAPD), simple sequence repeats (SSR), diversity array technology (DArT), and single nucleotide polymorphism (SNP) have been used to study genetic diversity and structure in crops (Kesawat and Das Kumar 2009). SSRs have the advantage to be abundant, highly polymorphic and multi-allelic, and therefore often provide more information compared to biallelic markers such as SNPs. On the other hand, the new iSelect genotyping platform, based on the Illumina Infinium assay, allows the simultaneous testing of 7842 gene-derived SNPs (Comadran et al. 2012). The genetic polymorphism of the SNP and SSR marker systems are generated through different mechanisms, thus they could give different views of the structure of a population. Analysing the genetic structure within a population is a critical step to the way to understand and reveal the complexity within this population (Pritchard et al. 2000). Factors such as human or environmentally driven selection, genetic drift, mating system and growth habit can have an effect on the population structure (Buckler and Thornsberry 2002; Flint-Garcia et al. 2003). Studies of worldwide (Malysheva-Otto et al. 2006), European (Rostoks et al. 2006), American (Hamblin et al. 2010) and Nordic (Rajala et al. 2016) barley germplasm have shown that cultivated barley has a clear level of population structure with major subpopulations caused by differences in ear type, i.e. two-row and six-row, and seasonal growth habit, i.e. winter and spring (Hamblin et al. 2010; MalyshevaOtto et al. 2006; Rostoks et al. 2006). In addition to these major subpopulations, it has been shown that American barley accessions further can be divided into minor sub-populations corresponding to the breeding programs, which might be allocated to limited exchange of material between the breeding programs and/or a result of local adaptation (Hamblin et al. 2010). Another important factor to consider is linkage disequilibrium (LD), which is the non-random association of alleles between two loci and shows the correlation between genetic polymorphisms, e.g. SNPs, and their history of mutations and recombination (Flint-Garcia et al. 2003). LD is important since the rate of its decay in a given species determines the number and density of the molecular markers needed to perform GWAS (Rafalski 2002). In many selfpollinated species such as barley where LD extends over long chromosomal distances (Malysheva-Otto et al. 2006), fewer markers are needed to cover the whole genome, whereas a higher marker density is needed when LD decays very rapidly in species such as maize where it declines to nominal levels within 1.5 kb (Remington et al. 2001). The aim of the Public Private Partnership (PPP) for pre-breeding in barley, partly funded by the Nordic Council of Ministers (NMR), is to lay a foundation for barley breeding for disease resistance and yield stability to meet current and future challenges in the Nordic region. This collaboration is between five breeding companies and three governmental organizations in the Nordic region. One of the goals with this program is to identify markers linked with traits of interest via genome-wide association studies (GWAS). The main objective with the present study is to determine the population structure and the LD decay in a Nordic barley panel, in order to estimate the relationships among individuals. Materials and methods Plant materials A total of 134 and 46 spring barley Hordeum vulgare L. subsp. vulgare breeding lines and cultivars, respectively, were included in this study. The selected spring barley breeding lines and cultivars are hereafter referred to as lines. Equal number of lines was selected Genet Resour Crop Evol 123 by breeders from each of Boreal Plant Breeding (Finland), Graminor Breeding AS (Norway), Agricultural University of Iceland (AUI Iceland), Lantma ¨nnen Lantbruk (LSW Sweden), Nordic Seed and Sejet Planteforaedling I/S (Denmark). The lines were chosen to represent the available genetic variation in current elite Nordic barley germplasm. Out of the 180 lines, eleven lines were removed from further analyses since they were duplicates, or due to incomplete genotyping. Out of the remaining 169 lines, 124 were two-rowed and 45 six-rowed. DNA extraction DNA was extracted from 2-week-old seedlings, using a CTAB (Cetyl Trimethyl Ammonium Bromide) method as described earlier by (Orabi et al. 2014). The DNA was precipitated with isopropanol, washed two times with 75% ethanol, air-dried and finally diluted in TE buffer (pH 8.0). Microsatellite genotyping All lines were genotyped using 48 microsatellite markers evenly distributed over all chromosomes. PCR amplifications were performed on a GeneAmp Ò PCR System 2700 thermal cycler (Applied Biosystems, Foster City, CA, USA) using a single universal touchdown PCR program as previously described in (Orabi et al. 2014). The forward primers were 50labeled with fluorescent dyes to achieve the maximum multiplex capacity of the ABI 3130xl sequencer. Direct and M13-labelling were used for the microsatellite fragments, with 6-carboxyfluorescein (6-FAM, blue) or hexachloro-6-carboxyfluorescein (HEX or VIC, green) and 50-fluorescein phosphoramidite (NED, yellow) for the direct labelling. For the M-13 labelling, 6-FAM (blue), VIC (green) and NED (yellow) were used. For fragment detection the ABI 3130xl DNA analyzer (Applied Biosystems, Foster City, CA, USA) was used and the fragment analysis and genotyping were performed using the GeneMarker genotyping software program, version 1.85 (Soft genetics, State College, PA, USA). SNP genotyping All lines were genotyped with the barley iSelect SNP chip based on the Illumina Infinium 9K assay. The genotyping of the lines was outsourced to Trait Genetics. The chip consists of 7842 high-confidence SNPs derived from expressed genes (Comadran et al. 2012). Data analysis Gene diversity and marker allele frequency Genetic distances between genotypes, genetic diversity, allele frequency and private alleles (alleles present only in one group) were calculated using an in-house program written in VBA (Visual Basic for Applications) and implemented in Microsoft Excel 2007 (Microsoft, Redmond, WA, USA). The program utilises R language software v.2.14.2 (R Development Core Team 2012), which includes the Modern Applied Statistics with S-plus (MASS) package (Ripley 2002). The average number of alleles per locus per group represents how polymorphic a given marker was within each group, and this value was calculated for each SNP marker. The average number of alleles per marker is between 1 and 2, where markers with a number of 1 were considered monomorphic. Modified Roger’s distances (MRD) based on (Wright 1978) were calculated for the SSR data based on the following equation: MRD ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 2mXm i¼1Xai k¼1ðpij qijÞ2 r where p ij and q ij are the allele frequencies of the jth allele at the ith marker of the two barley lines in consideration; a i is the number of alleles at the ith marker; and mis the number of SSR loci. The genetic distances (D SM ) based on the SNP data were determined by estimating a simple matching coefficient (S SM ) (Reif et al. 2005): DSM ¼1SSM;SSM ¼vij þyij vij þwij þxij þyij where v ij refers to the allele in common between two lines, iand j;w ij is the number of alleles present in i and absent in j;x ij is the number of alleles present in j and absent in i; and y ij is the number of alleles absent in both iand j. Correlations between the MRD (SSR) and D SM (SNP) matrices were calculated using Mantel test (Mantel 1967). The polymorphic information content (PIC) of the individual markers was calculated as explained by Botstein et al. (1980): Genet Resour Crop Evol 123 PIC ¼1X n i¼1 p2 iX n1 i¼1X n j¼iþ1 2p2 ip2 j where p i is the frequency of the ith allele, and nis the number of alleles per marker. To show genetic diversity and differentiation along the barley chromosomes Shannon’s diversity indices were calculated using GenAlEx v. 6.5.0.1 (Peakall and Smouse 2006,2012) based on the SNP data. Population structure analysis To determine population structure of the barley lines using SNP markers, the software package STRUCTURE v.2.3.4 based on a Bayesian clustering approach, was used (Pritchard et al. 2000). STRUCTURE was run 10 times for each hypothetical number of subpopulations (K) between 1 and 12 with the ploidy level set as 2. The Markov chain Monte Carlo (MCMC) was set to 9999 burn-in phases followed by 9999 iterations. Structure Harvester v.0.6.94 (Earl and von Holdt 2012), was used to estimate the most likely number of groups (K) using the DeltaK method (Evanno et al. 2005). Population structure based on the SSR markers was calculated as described above with the same settings but by using the previously described VBA program in Excel. Analysis of Molecular Variance (AMOVA), Nei’s unbiased genetic distance and Principal Coordinates Analysis (PCoA) were calculated using GenAlEx v. 6.5.0.1. Linkage disequilibrium analysis The TASSEL 3.0 software (http.//www. maizegenetics.net) was used to calculate the LD (allele frequency correlation, r 2 ) estimates between the SNP marker pairs using the full matrix option. Only intra-chromosomal comparisons were included and markers with minor allele frequency (MAF) below 0.05 were excluded. Thus 4884 out of the total 6280 polymorphic markers were subjected to analysis. To estimate the LD decay, the intra-chromosomal r 2 values were plotted against the genetic distance with a second-degree smoothed loess curve fitted using the program R (R Development Core Team 2012) and a baseline based on the critical value of r 2 was drawn. The critical value of r 2 , as an evidence of linkage, was calculated based on the method described in Breseghello and Sorrells (2006), by square root transforming the r 2 -values and taking the 95th percentile of unlinked r 2 -values. In the analysis markers located more than 50 cM apart were considered unlinked. Results Population structure in Nordic spring barley The STRUCTURE analysis indicated that the barley panel could be divided into two groups with common ancestry ([0.7) K1 SSR (n =109) and K2 SSR (n =50) and one group (n =10) with an admixed ancestry (\0.7), based on the SSR data. AMOVA analysis revealed that the K1 SSR ,K2 SSR and admixed SSR group were significantly separated (p\0.001) and explained 35% of total molecular variance (Online Resource 1). This is similar to the result obtained from the PCoA analysis, where the first two principal coordinates combined explained 32.1% (21.7 and 10.4%) of the variation (Fig. 1a). The subdivision of the population along the first principal coordinate (PC1) corresponded to the separation of the barley population into K1 SSR and K2 SSR . The K1 SSR and K2 SSR group corresponded to two-rowed lines and sixrowed lines, respectively, with the exception of six two-rowed lines found in K2 SSR . In the small admixed SSR group, nine two-rowed lines and one six-rowed line from mainly the northern parts were found. Also for the SNP data, two groups with a common ancestry ([0.7) (K1 SNP :n =109; K2 SNP :n =47) and one group (n =13) with an admixed ancestry (\0.7) were inferred by the STRUCTURE analysis. Three clusters were observed in the PCoA analysis, where the first and second principal components explained 24.6 and 5.9% of the variation, respectively, or 30.5% combined (Fig. 1b). AMOVA analysis revealed that the K1 SNP ,K2 SNP and admixed SNP groups were significantly separated (pvalue\0.001) and explained 42% of the total molecular variance (Online Resource 1). The two-rowed lines were distributed between K1 SNP and the admixed SNP group and the six-rowed lines were found in group K2 SNP , with the exception of two two-rowed lines that were found in the latter group. The lines with an admixed ancestry in the admixed SNP group were two-rowed and were, just as Genet Resour Crop Evol 123 the lines in the admixed SSR group, mainly from the northern parts. More detailed information regarding the origin of the lines and the inferred population structure groups can be found in Online Resource 2. Genetic diversity in Nordic spring barley Marker system comparison The barley lines were genotyped using 48 SSR markers. The same lines were also genotyped using the iSelect 9K SNP barley chip resulting in a total of 6208 polymorphic SNP markers. The SSR markers produced 234 scorable loci. The number of alleles per marker ranged from 1 to 15 with an average of 4.9 alleles per SSR marker. The average polymorphic information content (PIC) value was 0.46 for the SSR markers and 0.28 for the SNP markers. The genetic diversity index was higher for the SSR markers (0.514) than for the SNP markers (0.359). Mantel test showed a strong correlation between the SSR-based Modified Roger’s genetic distances and the SNP-based simple matching coefficient, with a Pearson’s value of r 2 =0.86 (pvalue\0.0000). Genetic diversity based on ear row type Allelic richness parameters in the form of average number of alleles per locus, number and proportion of private alleles and genetic diversity for each marker type based on ear row type are presented in Table 1. There were no major differences in the average number of alleles per locus found in the two-rowed lines (SSR, 4.0; SNP: 1.9) compared to the six-rowed lines (SSR: 3.5, SNP: 1.8).). The genetic diversity was higher in the two-rowed lines (SSR: 0.431; SNP: 0.305) compared to the six-rowed lines (SSR: 0.386; SNP: 0.225). The two-rowed lines had also a higher number of private alleles (SSR: 64; SNP: 1145) compared to the six-rowed lines (SSR: 41; SNP: 447). AMOVA analysis revealed that 35 and 40% of the molecular variance between the lines based on the SSRs and SNPs, respectively, could be explained by the row-types (data not shown). Genetic diversity based on population structure Allelic richness parameters in the form of average number of alleles per locus, number and proportion of private alleles and genetic diversity for each marker type based on the inferred population structure groups are presented in Table 2. For the SSRs the highest average number of alleles per locus and gene diversity was found in the admixed SSR group (3.7; 0.404). However, this group also had the lowest proportion of private alleles (0.30), whereas the highest proportion of private alleles was found within the six-rowed lines in group K2 SSR (0.86). No major differences in the average number of alleles per locus were found between the three subgroups, based on the SNP data. The highest proportion of private alleles (10.2) was, no different from the results obtained with the SSRs, found within the six-rowed lines in K2 SNP whereas the lowest Fig. 1 a Associations between structure groups revealed by principal coordinate analysis of the Nordic spring barley collection based on the SSR data. bAssociations between structure groups revealed by principal coordinate analysis of the Nordic spring barley collection based on the SNP data Genet Resour Crop Evol 123 proportion of private alleles (0.54) were found in the admixed SNP group. The highest gene diversity (0.279) was found in K1 SNP , whereas the lowest (0.198) was found in the admixed SNP group. Matrices showing relationships between the structure groups were generated based on Nei’s unbiased genetic distance for SSR and SNP data (Table 3). In the SSR matrix, the smallest distance (0.272) was found between the K1 SSR and admixed SSR group (both mainly two-rowed lines from the southern and northern parts, respectively). The largest distance (0.566) was found between groups K2 SSR (mainly six-rowed lines) and admixed SSR (mainly two-rowed lines from the southern and northern parts, respectively). Similar results were seen in the SNP matrix, where the smallest distance (0.217) was found between groups K1 SNP and admixed SNP (mainly two-rowed lines from the southern and northern parts, respectively) and the largest (0.338) between K2 SNP (mainly six-rowed lines) and admixed SNP (mainly two-rowed lines from the northern parts). When comparing genetic diversity along the barley chromosomes, differences between the population structure groups were evident in several genomic regions (Fig. 2). Both group K2 SNP and admixed SNP were low in diversity on chromosome 2H (around 83–113 cM), 3H (around 37–64 cM) and 5H (around 108–127 cM), whereas the K1 SNP lines were very diverse in these regions. A large difference in genetic diversity were also observed on chromosome 6H (around 55–58 cM), where the K1 SNP and K2 SNP lines were much more diverse compared to admixed SNP lines. Right before this region on chromosome 6H (around 49 cM), a region with low diversity was seen for the K1 SNP lines, but here the diversity was maintained for the two other groups. On chromosome 7H (around 68–88 cM) the lines in group K2 SNP were very diverse in contrast to the low diversity observed for the K1 SNP and admixed SNP lines in the same region. Linkage disequilibrium in Nordic spring barley LD for the whole population, the population structure groups and the ear row-types were calculated based on the SNP markers. The percentage of unlinked marker pairs ranged between 40 and 42% and no large differences were found between the different groups. The number of intra-chromosomal marker-pairs and the number of unlinked pairs in the total population and in the different groups are presented in Table 4.A total of 1,795,852 intra-chromosomal marker pairs were found in the entire population. The mean r 2 -value of the entire population, the ear row-types and population structure groups were calculated for the whole genome and for the seven chromosomes separately (Table 5). The mean r 2 -value for the whole genome of the entire population was found to be 0.10. The highest and lowest mean r 2 -value for the whole genome was found in the admixed SNP and K1 SNP group (r 2 : 0.19; 0.07), respectively. The interval in which the Loess curve intercepts the critical value (background LD) was considered as the LD decay. The background LD and the LD decay of the entire population, the ear row-types and population structure groups were calculated for the whole genome and for the seven chromosomes separately (Table 6, Online Resource 3). Average LD for the whole population decayed below the critical level (r 2 =0.20) within a range of 0–4.0 cM. However the LD decay for the different chromosomes, ear row-types and population structure groups ranged from 0 to 12, with the most extended LD decay found for 5H in the six-rowed lines and in group K1 SNP with an average r 2 of 8–12 Table 1 Allele frequency and genetic diversity in twoand six-rowed barley, based on SSR and SNP data Population size (n) Average number of alleles per locus Number of private alleles Proportion of private alleles Genetic diversity (D) Standard deviation (D) SSR SNP SSR SNP SSR SNP SSR SNP SSR SNP Two-row 124 4 1.9 64 1145 0.5 9.2 0.431 0.305 0.063 0.100 Six-row 45 3.5 1.8 41 447 0.9 9.9 0.386 0.225 0.123 0.233 Total 169 0.514 0.359 0.434 0.448 Genet Resour Crop Evol 123 (Table 6). When comparing the whole genome, the six-rowed lines, group K2 SNP and the lines from the northern parts in group admixed SNP had a more rapid LD decay with an average r 2 between 0 and 4 compared to the two-rowed lines and group K1 SNP , with lines from the southern parts, where the average LD decay were more slow with an average r 2 between 4–8 and 8–12, respectively. Discussion Comparison of marker systems It is of interest to compare if the information regarding population structure and genetic diversity is affected by the marker system of choice, since there are different mutational mechanisms behind SSR (replication slippage) and SNP (point mutation) markers. SSRs have been the most commonly used for studies of genetic diversity, mainly due to their abundance in the genome, reproducibility and high level of polymorphism. However, the increased availability of the SNP markers and the fast and highly automated genotyping technologies, have recently moved the attention to the use of the SNPs in studies of genetic diversity and population structure. This study showed that the average PIC and genetic diversity values were higher for the SSRs compared to the SNPs. However, these values have a maximum value of 0.5 for biallelic markers such as SNPs when the markers scores are 50% (0) and 50% (1). Taking this into Table 2 Private alleles and genetic diversity in the structure groups based on SSR and SNP data Group SSR data Group SNP data Population size Average number of alleles per locus No. of private alleles Proportion of private alleles Genetic Diversity (D) Standard deviation (D) Population size Average number of alleles per locus No of private alleles Proportion of private alleles Genetic Diversity (D) Standard deviation (D) K1 SSR 109 2.5 42 0.39 0.392 0.275 K1 SNP 109 1.9 639 5.86 0.279 0.153 K2 SSR 50 3.5 43 0.86 0.397 0.063 K2 SNP 47 1.8 481 10.23 0.235 0.208 Admixed SSR 10 3.7 3 0.30 0.404 0.120 Admixed SNP 13 1.6 7 0.54 0.198 0.495 Total 169 0.514 0.434 Total 169 0.359 0.448 Table 3 Nei’s unbiased genetic distance between different structure groups, based on the (a) SSR data, (b) SNP data K1 SSR K2 SSR Admixed SSR (a) K1 SSR 0.000 K2 SSR 0.292 0.000 Admixed SSR 0.272 0.566 0.000 K1 SNP K2 SNP Admixed SNP (b) K1 SNP 0.000 K2 SNP 0.250 0.000 Admixed SNP 0.217 0.338 0.000 Genet Resour Crop Evol 123 consideration, the SNPs would be just as or even more informative than the SSRs. A higher number of private alleles were found with the SNPs compared to the SSRs (Tables 1,2). However, considering the different number of markers, the SSRs actually had the highest private allele frequency, which could be expected since the SNPs are bi-allelic and have a lower mutation rate compared to the SSRs (MartinezArias et al. 2001; Li et al. 1981; Kruglyak et al. 1998). It has earlier been reported that there is strong correlation between these two marker systems in barley (Varshney et al. 2008; Varshney et al. 2010). Fig. 2 Shannon’s diversity index calculated as rolling means over 20 adjacent loci. The start and end position of each chromosome are indicated with vertical lines at the bottom of the figure Table 4 Number of intrachromosomal marker-pairs in the total population and in the different groups Unlinked marker-pairs refers to a marker-pair distance [50 cM Population size Total pairs Unlinked pairs Unlinked pairs (%) Total population 169 1,795,852 744,332 41 Two-row lines 124 1,464,051 609,353 42 Six-row lines 45 923,896 367,226 40 K1 SNP 109 1,166,044 482,154 41 K2 SNP 47 981,129 393,735 40 Admixed SNP 13 663,021 262,178 40 Table 5 Mean r 2 -values for intra-chromosomal marker-pairs in the whole genome and for each chromosome Chromosome No. of SNPs Total population Two-rowed lines (n =124) Six-rowed lines (n =45) K1 SNP (n =109) K2 SNP (n =47) Admixed SNP (n =13) 1H 437 0.12 0.05 0.11 0.05 0.11 0.15 2H 768 0.11 0.07 0.12 0.06 0.11 0.23 3H 724 0.07 0.08 0.10 0.08 0.10 0.17 4H 578 0.09 0.07 0.11 0.06 0.10 0.18 5H 1025 0.10 0.10 0.10 0.05 0.10 0.21 6H 726 0.09 0.09 0.14 0.10 0.13 0.17 7H 626 0.09 0.07 0.08 0.07 0.09 0.19 Whole genome 4884 0.10 0.08 0.11 0.07 0.10 0.19 Genet Resour Crop Evol 123 This was also demonstrated here with a strong and positive correlation (r =0.86) between the modified Roger’s distances based on the SSR data and the simple matching coefficient based on the SNP data. Also the PCoA and structure analysis grouped the lines in a similar way with the two marker systems. However, the PCoA clustering based on the SNP data showed clearer and more distinct groupings than the SSR-based PCoA, but no major difference in the amount of molecular variance explained was found between the two marker systems. Neither did the AMOVA analyses of the structure groups reveal any large differences between the two marker systems (Table S1). However, superimposing the results from the population structure analysis on the results from the PCoA provided a clearer image and higher resolution of the population structure based on the SNPs compared to the SSRs (Fig. 1a, b). That reveals the ability of SNPs to explain the population structure at a more specific level. This was also seen in a study of genetic diversity and population structure of 375 rice varieties (Singh et al. 2013), where a comparison of the SSR and SNP marker systems revealed that at the structure level the SNPs were better at describing genetic relatedness whereas at the diversity level the SSRs showed a better grouping of samples. In contrast, some studies of genetic diversity and population structure in maize report a better estimate of population structure with SSRs compared to SNPs (Yang et al. 2011; Hamblin et al. 2007). The different results between those reports and this study might be due to the different numbers of SNP markers used (\900 vs. 6208) or due to the complexity of the maize genome. According to a theoretical prediction by Laval et al. (2002), (k-1) times more bi-allelic markers are needed to achieve a comparable accuracy of the genetic distance as a set of SSRs with kalleles. With the average of about 3 alleles per SSR marker in this study, the number of SNPs needed would be [(3 -1) 948] =96, which are about 65 times less compared to the 6208 used. Altogether the results from this study show that the numbers of SNPs used are more than enough to retrieve a comparable accuracy of genetic diversity in barley as the set of SSRs used. In addition, the SNPs seem to provide a higher resolution for the genetic relatedness than obtained with the SSRs. Diversity and relationships within Nordic spring barley The average genetic diversity for the Nordic spring barley collection analysed here was 0.514 and 0.359 based on the SSRs and SNPs, respectively. These are similar to the diversity estimates of Nordic breeding lines and cultivars released after 1970 (0.601) reported in an earlier study based on SSRs (Kolodinska Brantestam et al. 2007). In addition a similar result based on SSRs was reported in accessions from Europe (0.593), Eritrea (0.573) and Ethiopia (0.620), whereas the Hordeum vulgare subsp. spontaneum (K. Koch) and H. vulgare accessions from the West Asia North Africa (WANA) region had a higher diversity (0.826 and 0.762, respectively) (Orabi et al. 2007). Table 6 Interval of the estimated LD decay (cM) in the total population and for the different groups Chromosome No. of SNPs Total population Two-rowed lines (n =124) Six-rowed lines (n =45) K1 SNP (n =109) K2 SNP (n =47) Admixed SNP (n =13) 1H 437 0–3 6–9 3–6 6–9 3–6 3–6 2H 768 0–3 3–6 3–6 6–9 3–6 0–3 3H 724 7–11 7–11 4–7 7–11 4–7 4–7 4H 578 0–3 3–5 0–3 3–5 0–3 0–3 5H 1025 0–4 0–4 0–4 8–12 4–8 0–4 6H 726 0–3 5–8 0–3 5–8 0–3 0–3 7H 626 0–3 7–10 7–10 7–10 3–7 0–3 Whole genome 4884 0–4 4–8 0–4 8–12 0–4 0–4 Background LD whole genome a 0.20 0.16 0.19 0.10 0.19 0.31 a The 95th percentile of unlinked (above 50 cM) square root transformed r 2 values Genet Resour Crop Evol 123