scieee AI-readable full text Open interactive document viewer

Construction of a map-based reference genome sequence for barley, Hordeum vulgare L.

Beier, S.,Himmelbach, A.,Colmsee, Chr.,Zhang, X.-Q.,Barrero, R. A.,Zhang, Q.,Li, L.,Bayer, M.,Bolser, D.,Taudien, St.,Groth, M.,Felder, M.,Hastie, A.,Simkova, H.,Stankova, H.,Vrana, J.,Chan, S.,Munoz-Amatriain, M.,Ounit, R.,Wanamaker, St.,Schmutzer, Th.,

Full text

Data Descriptor: Construction of a map-based reference genome sequence for barley, Hordeum vulgare L. Sebastian Beier et al. # Barley (Hordeum vulgare L.) is a cereal grass mainly used as animal fodder and raw material for the malting industry. The map-based reference genome sequence of barley cv. ‘Morex’was constructed by the International Barley Genome Sequencing Consortium (IBSC) using hierarchical shotgun sequencing. Here, we report the experimental and computational procedures to (i) sequence and assemble more than 80,000 bacterial artificial chromosome (BAC) clones along the minimum tiling path of a genome-wide physical map, (ii) find and validate overlaps between adjacent BACs, (iii) construct 4,265 non-redundant sequence scaffolds representing clusters of overlapping BACs, and (iv) order and orient these BAC clusters along the seven barley chromosomes using positional information provided by dense genetic maps, an optical map and chromosome conformation capture sequencing (Hi-C). Integrative access to these sequence and mapping resources is provided by the barley genome explorer (BARLEX). Design Type(s) genome assembly Measurement Type(s) whole genome sequencing assay Technology Type(s) DNA sequencing Factor Type(s) library preparation Sample Characteristic(s) Hordeum vulgare Correspondence and requests for materials should be addressed to M.M. (email: [email protected]). #A full list of authors and their affiliations appears at the end of the paper. OPEN Received: 26 August 2016 Accepted: 9February 2017 Published: 27 April 2017 www.nature.com/scientificdata SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 1 Background & Summary Barley (Hordeum vulgare L.) is a cereal grass of great agronomical importance. The goal of the International Barley Genome Sequencing Consortium (IBSC) is the construction of a map-based reference sequence assembly of barley cultivar ‘Morex’by means of hierarchical shotgun sequencing 1 . Towards this aim, the barley genomics community has developed an array of genome-wide physical and genetic mapping resources. These include libraries of bacterial artificial chromosomes (BACs) 2 , a genome-wide physical map 3 , a draft whole genome shotgun (WGS) assembly 4 and an ultra-dense genetic map 5 . The last stage on the road towards the reference genome is the shotgun sequencing of BAC clones along a minimum tiling path of the genome defined by the physical map. The advances in highthroughput sequencing technology enabled this task to be completed in a much shorter timeframe than was required for the completion of, for instance, the human 6 and maize 7 genomes. In addition to the generation of BAC raw sequence data, we constructed (i) physical genome maps by single-molecule optical mapping in nanochannels 8 and by chromosome conformation capture sequencing (Hi-C) 9,10 , and (ii) a high-resolution genetic map of a large bi-parental mapping population through genotyping-bysequencing 11 . We undertook the sequence assembly of individual BACs, the construction of larger sequence scaffolds by merging sequences from adjacent clones and the integration of these superscaffolds with the various genome-wide mapping resources constructed in the present effort as well as those published previously 3,5 . The final outcome of this approach was the construction of ‘pseudomolecules’, i.e., contiguous sequence scaffolds representing the seven chromosomes of barley. We have submitted the relevant raw data to public sequence data archives, made analysis results available under permanent digital object identifiers (DOIs) and entered the positional information used for pseudomolecule construction into a bespoke information management system, the BARLEX genome explorer 12 . Here, we give (i) a comprehensive overview of datasets used for assembling the barley genome and methods employed in their generation, (ii) a detailed description of wet-lab procedures for BAC sequencing and the bioinformatics workflow of the sequence assembly and data integration procedures together with an outline of (iii) their browsable presentation in an online database. These resources document the construction of the map-based reference sequence of the barley genome and will enable researchers to inspect the evidence used to assemble, order and orient sequence scaffolds and may guide the further improvement of the genome sequence with complementary data sets. Methods The main steps for the construction of the map-based reference sequence of the barley genome were (i) shotgun and mate-pair sequencing of BAC clones, (ii) sequence assembly of individual BAC clones and (iii) the construction of a pseudomolecule sequences by merging the sequences of adjacent BACs into super-scaffolds and ordering these using various sources of positional information such as physical maps, optical map and chromosome conformation capture. A schematic overview of our experimental procedures is given in Fig. 1. BAC sequencing Identification and analysis of gene-containing BACs. Isolation of gene-containing BACs, construction of a minimal tiling path (MTP), sequencing of MTP clones and the annotation of genes were essentially as described previously 13 . Shotgun and mate-pair sequencing of MTP-BACs. Sequencing of MTP-BACs was conducted in four laboratories (Leibniz Institute on Aging—Fritz Lipmann Institute (FLI) Jena, Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Beijing Genomics Institute (BGI) and Earlham Institute (EI) Norwich). Depending on the instrumentation and established protocols, customized approaches were taken to sequence the barley MTP BACs. Barley chromosomes 1H, 3H and 4H (IPK and FLI) Shotgun sequencing of MTP BACs During the initial phase, BACs mostly from chromosome 3H (4870 clones) and a small number of clones from other chromosomes (34 from 1H; 31 from 2H; 50 from 4H; 101 from 5H; 33 from 6H; 64 from 7H; 107 from ‘0H’) were shotgun sequenced using the Roche/454 GS FLX device (Data Citation 1, Data Citation 2, Data Citation 3, Data Citation 4, Data Citation 5, Data Citation 6, Data Citation 7, Data Citation 8, Data Citation 9). BAC DNA was prepared using a modified alkaline lysis protocol 14 . Construction of barcoded 454 sequencing libraries and sequencing using the Roche platform were performed as described 15,16 . The remaining BAC clones from chromosomes 1H, 3H and 4H were shotgun sequenced employing Illumina instruments. BAC DNA isolation, library construction, sequencing-by-synthesis (paired-end, 2 × 100 cycles) using the Illumina HiSeq2000 device was performed as described 17 (Data Citation 10, Data Citation 11, Data Citation 12, Data Citation 13). Pools of up to 667 BACs were individually barcoded and sequenced on one HiSeq2000 lane. In addition, the Illumina GAIIx, HiSeq2500 and MiSeq machines were utilized to sequence pools of up to 384 clones per lane as described previously 17 . www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 2 Mate-pair sequencing of MTP BACs For scaffolding of chromosomes 1H, 3H and 4H standard Illumina Nextera mate-pair libraries (span size: 8 kb) of BAC pools up to 384 BACs were constructed and sequenced using the Illumina HiSeq2000 (paired end, 2 × 100 cycles) and MiSeq (paired end, 2 × 250 cycles) as described 17 (Data Citation 14, Data Citation 15). Barley chromosomes 5H, 6H and 7H (BGI) Shotgun sequencing of MTP BACs Bacterial starter cultures were inoculated in 0.4 ml 2 × YT liquid medium 18 supplemented with chloramphenicol (17.5 μgml 1 ) in 2 ml polypropylene 96-deep well-plates sealed with gas-permeable foil and incubated at 37 °C for 14 h in a shaking incubator (210 r.p.m.). For DNA isolation duplicates of cultures (1 ml 2 × YT liquid medium containing 17.5 μgml 1 chloramphenicol) were inoculated with 50 μl starter culture and incubated (37 °C, 14 h, 210 r.p.m.). BAC DNA was isolated using the alkaline lysis method essentially as described previously 17 . The DNA was dissolved (overnight, 4 °C) in 64 μlTE (pH 8.0) containing RNase A (30 μgml 1 ) and stored at −20 °C. BAC plasmid DNA (0.5–2.0 μgin60μl) was randomly fragmented by focused-ultrasonicator (Covaris LE220 instrument: 21% duty factor, 500 PIP, 500 cycles per burst, 70 s treatment time) in 96-well plates (Axygen, PCR-96M2-HS-C) to an average size of 250–750 bp. The DNA fragments were purified using magnetic beads (GeneOn Purification kit, GO-PCRC-5000) according to the manufacturer’s instructions. DNA was precipitated by adding 10 μl magnetic bead suspension and 75 μl Binding Buffer. The samples were mixed and incubated at room temperature for 5 min. Beads containing the DNA were reclaimed by using a magnet (96S Super Magnet Plate, ALPAQUA, A001322), and the clear supernatant was discarded. The beads were washed twice with 200 μl of 70% ethanol and dried completely. For the elution of DNA the beads were suspended in 42 μl Elution Buffer (EB, 10 mM Tris-Cl, pH 8.5) and incubated (5 min). The plate was placed on the magnet, and the supernatant (40 μl) was transferred into new 96-well plates. End-repair and A-Tailing were performed as described 19 . The reaction clean-ups were performed with GeneOn magnetic beads as described above. Barcode adapters (1 μl, 20 μM) for the first index were ligated to the sticky ends of DNA fragments by using T4 DNA ligase 19 , incubated at 16 °C for at least 12 h. Each individual sample was provided with a different barcode of a set of 384 different indices (adapter and barcode sequences are available upon request). Equal volumes of the 384 individually barcoded adapter-ligated products were pooled. The pooled DNA was precipitated by adding 20 μl GeneOn magnetic beads and 650 μl Binding Buffer (GeneOn Purification Kit, GO-PCRC-5000) BAC DNA preparation Paired-end library construction Mate-pair library construction Quantification, pooling and size fractionation Sequencing-bysynthesis (Illumina) Removal of low quality sequences & contamination Individual BAC assembly Quantification, pooling and size fractionation Sequencing-bysynthesis (Illumina) Removal of low quality sequences & contamination Mapping & individual BAC scaffolding Individual BAC scaffolds BAC scaffolds FPC / BES data POPSEQ map Bionano map data Conformation capture data (HiC / TCC) BAC overlap clusters BLAST analysis Nonredundant sequence Conformation capture map (HiC map) AGP generation & Pseudomolecule sequence Figure 1. Assembly workflow. (a) Assembly of individual BAC clones from paired-end and mate-pair read data. (b) Data integration procedures for pseudomolecule construction. www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 3 to 500 μl pooled DNA. The suspension was mixed and incubated at room temperature for 5 min. The beads containing the DNA were reclaimed using a magnet, and the clear supernatant was discarded. The beads were washed twice with 500 μl of 70% ethanol and dried completely. The DNA was eluted in 52 μl EB. The sample was size-separated by using standard agarose gel electrophoresis (2% agarose gel, HyAgarose, 16250). DNA was revealed using ethidium bromide and excitation by visible blue light emitted from a Dark Reader blue light transilluminator (Clare Chemical Research) to select the target fragments (580–620 bp). The target region was extracted in 27 μl EB using the QIAquick Gel Extraction kit (QIAGEN). The second index was introduced using the adapter-ligated products as template DNA (98 °C for 30 s, 10 cycles of: 98 °C for 10 s, 65 °C for 30 s and 72 °C for 30 s, final extension 72 °C for 5 min) (Enzymatics, CM0075) and PCR products (target region: 580–620 bp) were recovered by agarose gel electrophoresis (2% agarose gel, HyAgarose, 16250) as described above. Index primers were used for barcoding each 384 pooled BAC samples (index primer sequences are available upon request). The average size of the PCR products was determined by using an Agilent 2100 Bioanalyzer (Agilent DNA 1,000 Reagents). Typical average size of the libraries was between 574 to 674 bp. PCR products were quantified using real-time PCR and pooled for sequencing in equal proportion 20 . Paired-end sequencing (2 × 100 cycles; first index: 11 cycles, second index: 8 cycles) was performed on the Illumina HiSeq2000 platform (Data Citation 16, Data Citation 17, Data Citation 18). Mate-pair sequencing of MTP BACs For the construction of mate-pair libraries (10 and 20 kb span size), 96 BACs corresponding to 6 μg DNA were pooled into one tube. The DNA was fragmented to 10 or 20 kb by using the HydroShear DNA Shearing system from GeneMachines (10 kb: large assembly, speed code 12, cycles 12, volume 250 μl; 20 kb: large assembly, speed code 13, cycles 20, volume 150 μl). Following DNA fragmentation, the fragments were purified by using 0.6 volumes magnetic beads (Axygen, MAG-PCR-CL-250). The samples were mixed and incubated at room temperature for 10 min. Beads containing the DNA were reclaimed by using a magnet plate (96S Super Magnet Plate, ALPAQUA, A001322), and the clear supernatant was discarded. The beads were washed twice with 500 μl of 70% ethanol and dried completely. For the elution of DNA the beads were resuspended in 80 μl EB. End-repair and biotin-labeling were performed as described 21 . End-repaired DNA was purified using 0.6 volumes magnetic beads (Axygen, MAG-PCRCL-250) as described for the purification of hydro-sheared DNA. The DNA was eluted in 79 μl EB. 20 kb libraries (20–26 kb range) were size-selected using agarose gel (0.6%) electrophoresis. The ligation of the libraries, was performed by adding 1 μl Barcode Adaptor (20 μM, sequences are available upon request), 10 μl T4 DNA ligase (Enzymatics, L603-HC) in a total volume of 100 μl (20 °C, 15 min). 15 individually barcoded adaptor-ligated DNAs (10 kb) were pooled in equimolar manner and size-fractionated (9–11 kb) using agarose gel (0.6%) electrophoresis. DNA circularization and removal of non-circularized DNA was as described 21 . The DNA was isolated from the gel using the QIAquick Gel Extraction kit as described by the manufacturer (QIAGEN). Circular DNA was fragmented using the Covaris S2 device (10% duty cycle, 10 intensity, 1,000 bursts per second, 22 min (11 min) treatment time for 10 kb (20 kb) libraries in TC13 Covaris tubes), and biotinylated fragments derived from true mate-pair ligation events were purified using streptavidin-coupled Dynabeads (M-280, Invitrogen) 19 . Ends of the DNA fragments were repaired and provided with Illumina paired-end adapters as described for the construction of shotgun libraries. The bead-bound DNA was PCR-amplified using Phusion polymerase (NEB) (98 °C for 30 s, 18 cycles of: 98 °C for 10 s, 65 °C for 30 s, 72 °C for 30 s and a final extension: 72 °C for 5 min) using manufacturer’s protocols (NEB). Size-selection was essentially performed as described for shotgun library construction. For the 10 kb (20 kb) mate-pair libraries, DNA in the size range between 270–420 bp (400–600 bp) was isolated and purified using the QIAquick Gel Extraction kit according to manufacturer’s instructions (QIAGEN). The average size of the paired-end BAC libraries was determined electrophoretically using an Agilent 2100 Bioanalyzer (Agilent DNA 1,000 Reagents). Libraries were quantified using Real-Time PCR 20 . The mate-pair libraries were paired-end sequenced using the Illumina HiSeq2500 device (10 kb library: 150 cycles, 20 kb mate-pair library 50 cycles). Raw data are available as Data Citation 19, Data Citation 20, Data Citation 21). Barley chromosomes 2H and 0H (EI) Shotgun sequencing of MTP BACs QRep 384 Pin Replicators (Molecular Devices, New Molton, UK) were used to inoculate clones from stock plates into 384 square deep well culture plates containing 140 μl 2 × YT media supplemented with 12.5 μgml 1 chloramphenicol 18 . The culture plates were sealed with a gas permeable seal and incubated for 22 h at 37 °C in a shaking incubator (200 r.p.m.). Cells were harvested by centrifugation (20 min, 3,220 g, 4 °C), the supernatant was discarded. BAC DNA was prepared using a modified alkaline lysis protocol (Beckman Coultier, High Wycombe, UK). Cell pellets were resuspended in 8 μl of Resuspension Buffer (RE1) using a Microplate Shaker TiMix 5 control (Edmund-Buehler, Hechingen, Germany) (10 min, 1,400 r.p.m.). Cells were lysed by adding 8 μl of the lysis solution (L2). After shaking (5 min, 500 r.p.m.) 8 μl of cold Neutralisation Buffer (N3) were added. The plate was shaken (10 min, 500 r.p.m.) followed by a centrifugation (20 min, 3,220 g, 4 °C). The clear supernatant (14.33 μl) was transferred www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 4 to a 384 well PCR plate, which contained 1 μl of CosMc beads per well. The plate was mixed briefly (500 r.p.m.), 10 μl of isopropanol was added and the suspension was mixed briefly again (500 r.p.m.). The plate was incubated at room temperature for 15 min to allow precipitation of the DNA onto the beads. The plate containing the DNA precipitate was moved onto a 96 pin 384 well plate compatible magnet (Alpaqua, Beverley, MA, USA) and left for 5 min for the beads to pellet. The supernatant was discarded and the beads were washed three times with 20 μl 70% ethanol while placed in the magnet and air dried (room temperature, 5 min). The DNA was eluted from the beads in 20 μl of 10 mM Tris HCl (pH 8.0) and transferred to a fresh 384 well PCR plate. To remove contaminating host E. coli gDNA samples were treated with Epicentre Plasmid Safe ATP dependent DNase (Cambio, Cambridge, UK), which digests the fragmented E. coli and nicked BAC DNA but leaves supercoiled BAC DNA intact. To 20 μl of DNA 2.5 μl of 10x Reaction buffer, 1 μl 25 mM ATP, 0.1 μl ATP dependent DNase (10 u μl 1 ) and 1.4 μl water was added, and the samples were incubated at 37 °C (8 h) followed by 70 °C (20 min) to inactivate the DNase. Sequencing libraries (single index) from the initial sixteen 384 well plates of BACs (2H chromosome) were constructed in 384 well PCR plates (Fortitude, Wotton, UK) using the Epicentre Nextera Kit (Epicentre, Madison, WI, USA) and Robust 2G Taq polymerase (Kapa Biosciences, London, UK). The 384 adapter oligos with 9 bp barcodes each with a hamming distance of 4 (adapter sequences are available upon request) were designed using standard guidelines 22 . Briefly, 1 μl of BAC DNA, 1 μl Nextera HMW 5 × Reaction Buffer, 1 μl of Nextera Enzyme (diluted 50-fold in 50% glycerol, 0.5 × TE pH 8.0) and 2 μlof water were combined and incubated (5 min, 55 °C) as described 23 . For the denaturation of the Tn5 polymerase, 15 μl PB Buffer (Qiagen, Manchester, UK) and for the reaction clean-up, 20 μl AMPure XP (Beckman, High Wycombe, UK) beads were added using a Caliper Sciclone Robot (Perkin Elmer, Coventry, UK). Following an incubation (5 min, room temperature), the precipitated tagmented DNA was purified using a 96 well ring Magnet (Alpaqua, Beverly, MA, USA). The beads were washed twice with 20 μl 70% ethanol while placed in the magnet before being air dried for 5 min. The tagmented DNA was eluted in 5 μl 10 mM Tris HCl, pH 8.0 and transferred to a fresh 384 well PCR plate. To 5 μl purified, tagmented DNA 2 μl of 5 × 2G B Reaction buffer, 0.2 μl of 10 mM dNTPs, 0.1 μl of Robust 2G Taq polymerase, 0.2 μl of 50 × Nextera Primer Cocktail and 2.5 μl 0.2 μM barcoded P2 adapter primer were added in a total reaction volume of 10 μl and amplified according to the following thermal cycling profile: 72 °C for 3 min, 95 °C for 1 min, followed by 21 cycles of 95 °C for 10 s, 65 °C for 20 s and 72 °C for 3 min. Post amplification the DNA concentration was determined using the Quant-It Picogreen dsDNA assay (Thermo Fisher, Cambridge, UK). Library DNA concentrations typically ranged from 4 to 40 ng μl 1 (average of 16 ng μl 1 ). For each sample from a 384 well plate a 5 μl aliquot was pooled and split into two 2 ml Lo bind Eppendorf tubes (950 μl each). To each aliquot 950 μl of AMPure XP (Beckman, High Wycombe, UK) beads was added. Samples were mixed, incubated (5 min, room temperature) and placed on a magnet particle concentrator (MPC) until the beads were collected. The supernatant was discarded. The beads were washed twice with 20 μl 70% ethanol while placed in the MPC and air dried (5 min). The pooled library was eluted from the beads in 17 μl of 10 mM Tris HCl pH 8.0. The two 17 μl aliquots of the library were combined and the DNA concentration was determined using the Qbit device with the Quant-It DNA HS Assay (Invitrogen). Typical DNA concentrations were above 100 ng μl 1 . The DNA size selection was performed using the Blue Pippin (Sage Science, Beverly, MA, USA). About 3 μg of the library in 30 μl of 10 mM Tris HCl pH 8.0 and 10 μl of the R2 ladder were separated (tight selection protocol, 650 bp) using a 1.5% agarose cassette according to the manufacturer’s instructions (Sage Science, Beverly, MA, USA), thereby yielding an average insert size of about 485 bp. Size selected samples were collected in 40 μl of TRISTAPS buffer, pH 8.0 (Sage Science, Beverly, MA, USA). The average size of the library was determined using a High Sensitivity Chip and an Agilent 2100 Electrophoresis Bioanalyzer (Agilent). The DNA concentration was measured using the Qbit device and the Quant-It DNA HS Assay (Invitrogen). Size selected libraries were quantified using the Kappa Biosciences Illumina library qPCR quantification kit (Kapa Biosciences) on a Step One qPCR machine (ThermoFisher) according to the manufacturer’s instructions and compared against a known concentration of a PhiX control library. Several libraries were pooled for sequencing in an equimolar manner, and the final pool was re-quantified for sequencing relative to a standard library of a known concentration using the Kapa Biosciences Illumina library qPCR quantification kit. Sequencingby-synthesis for 6,144 BACs from chromosome 2H was performed using an Illumina HiSeq2000 device (2 × 100 cycles paired-end, single indexing read, 384 BACs/lane) according to manufacturer’s instructions, thereby yielding at least 32 Gb/lane and an average sequence coverage of at least 500-fold per BAC. The remaining BAC clones from 2H (384 BACs/lane) and 0H (2304 BACs/lane) were sequenced with a HiSeq2500 machine (2 × 150 cycles paired-end, dual indexing, rapid mode, yield: at least 30 Gb/lane) using a slightly adapted protocol with an additional normalization step prior to sample pooling. Briefly, a custom panel of 48 P5 and 48 P7 adapter oligos with 9 bp barcodes (with ≥4 hamming distance) was designed to individually label up to 2,304 (48 × 48) libraries by dual indexing. A mixture of 2μl of BAC DNA, 0.5 μl Nextera 10 × Reaction Buffer, 0.1 μl Nextera Enzyme and 2.4 μl water was incubated (5 min, 55 °C). Tn5 denaturation, reaction clean-up, washing, elution and transfer to a fresh 384 well plate were as described for the single-indexing libraries. 5 μl purified, tagmented DNA, 2 μlof 5 × Kapa Robust 2G B Reaction buffer, 0.2 μl of 10 mM dNTPs, 0.05 μl of Kapa Robust 2G Taq polymerase, 1 μl2μM P5 primer, 1 μl2μM P7 primer were combined (reaction volume of 10 μl) and amplified according to following thermal cycling profile: 72 °C for 3 min, 95 °C for 1 min, followed by www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 5 16 cycles of 95 °C for 10 s, 65 °C for 20 s and 72 °C for 3 min. The size profile and quantity was determined as described for single-indexing libraries. Amplified libraries were normalised using MagQuant bead technology (GC Biotech, Netherlands) on a Caliper Zephyr Robot (Perkin Elmer), essentially as described by the manufacturer. Normalised libraries were eluted in 10 μl of 10 mM Tris HCl pH 8.0 and transferred to a fresh 384 well PCR plate.5 μl of 384 normalized samples were pooled (total volume 1,920 μl). Purification using AMPure XP beads, washing, elution, size-selection (Blue Pippin) and quality checks prior to sequencing were essentially as described for single indexing libraries. Sequencing-by-synthesis of pooled libraries (2,304 BACs) was performed using an Illumina HiSeq2500 device (rapid run mode, 2 × 150 cycles paired-end, dual indexing reads) according to manufacturer’s instructions. At least 40 Gbp/lane, and an average sequence coverage of >100-fold per BAC were obtained (Data Citation 22, Data Citation 23, Data Citation 24, Data Citation 25). Mate-pair sequencing of MTP BACs BAC clones were inoculated as described for the preparation of shotgun libraries. The bacterial cultures were grown for 6 h at 37 °C in a shaking incubator at 200 r.p.m., and 384 clones were pooled. The pool was used to inoculate 250 ml 2 × YT media supplemented with chloramphenicol (12.5 μgml 1 ). The cultures were incubated (18 h, 37 °C, 200 r.p.m.). Cells were harvested by centrifugation (3,220 g, 20 min, 4 °C), and the supernatant was discarded. Alkali lysis and DNA isolation steps were performed using the Large Construct kit (Qiagen, UK) essentially following the manufacturer’s instructions. The DNA was resuspended in 4.75 ml Buffer Ex, 100 μl 100 mM ATP (Fisher Scientific, UK) were added and contaminating E. coli DNA was removed using 150 μl ATP dependent Exonuclease (Qiagen). During the incubation (1 h, 37 °C) a Qiagen Tip-100 column (Qiagen) was equilibrated in Buffer QBT (Qiagen). 5 ml of Buffer QS were added to the DNA, and the sample was applied to the equilibrated column. The column was washed twice with 10 ml of Buffer QC (Qiagen). The DNA was eluted with 7.5 ml of pre-warmed (65 °C) Buffer QF (Qiagen). The DNA was precipitated by adding 0.7 × volume of room temperature isopropanol and centrifugation (20 min, 3,220 g, 4 °C). The pellet was washed twice with 70% ethanol, air dried and dissolved in 200 μl TE buffer according to manufacturer’s guidelines. The DNA concentration was measured using a Qubit Fluorometer (Thermo Fisher, Cambridge, UK) and adjusted with water to 13 ng μl 1 . For tagmentation 200 μl diluted DNA were equilibrated (6 min, 55 °C) and subsequently provided with 52 μl 5 × Tagment Buffer Mate-Pair and 8 μl Mate-Pair Tagmentation Enzyme (Illumina, San Diego, USA). After the incubation (30 min, 55 °C), 65 μl Neutralize Tagment Buffer (Illumina, San Diego, USA) were added, and the reaction was incubated (5 min, room temperature). One volume CleanPCR beads (GC Biotech, Alphen aan den Rijn, The Netherlands) was added, and the DNA was purified using magnetic separation. The DNA was eluted in 170 μl of nucleasefree water, quantified using a Qubit fluorometer (DNA HS assay, Invitrogen) and analysed using the Agilent Bioanalyser (DNA 1,200 chip, Agilent, Stockport, UK). Strand displacement was performed by combining 105.3 μl of tagmented DNA, 13 μl 10x Strand Displacement Buffer (Illumina), 5.2 μl dNTPs (Illumina), 6.5 μl Strand Displacement Polymerase (Illumina) and incubation (30 min, room temperature). CleanPCR beads (0.75 volume) were added and the DNA was purified using a magnet. The DNA was eluted in 30 μl nuclease-free water. The concentration was measured (Qubit, DNA HS assay, Invitrogen), and a 1:6 diluted sample was analysed using the Agilent Bioanalyser (DNA 1,200 chip, Agilent, Stockport, UK). Size selection was performed using a Pippin Blue (Sage Science, Beverly, MA, USA). 30 μl DNA were provided with 10 μl loading buffer and separated on a 0.75% agarose cassette (size selection centered at 7 kb and collection between 6–8 kb) according to the manufacturer’s instructions (Sage Science, Beverly, MA, USA). Size selected samples were collected in 40 μlof TRISTAPS buffer (pH 8.0) (Sage Science, Beverly, MA, USA), and analysed using the Agilent Bioanalyser (high sensitivity chip, Agilent, Stockport, UK) to determine the final library size. The DNA concentration was measured using the Qubit device and the Quant-It DNA HS Assay (Invitrogen). Circularisation was performed by combining 40 μl size selected DNA, 12.5 μl 10 × circularisation buffer (Illumina), 3 μl Circularisation Enzyme (Illumina) and 75 μl nuclease-free water. The reaction was incubated at 30 °C overnight. Linear DNA was digested by adding 3.75 μl Exonuclease (Illumina) and incubation (30 min, 37 °C). The enzyme was inactivated by heat (30 min, 70 °C) and the addition of 5 μl stop ligation (Illumina). Circularised DNA (130 μl) was sheared in a Covaris MicroTube AFA Fiber (Pre-slit, Snap-cap, 6 × 16 mm; 2 cycles of 37 s, 10% duty cycle, 200 cycles per burst, 4 intensity, 4 °C) using the Covaris S2 device (Covaris, Massachusetts, USA). M280 Dynabeads (Thermo Fisher) were prepared as described (Illumina). 130 μl washed M280 beads were added to the fragmented DNA, mixed and placed on a lab rotator (20 min, room temperature). Library molecules were affinity purified and washed as described (Illumina). The beads were resuspended in a mixture of 85 μl nuclease free water, 10 μl 10x End Repair Reaction Buffer (Ilumina) and 5 μl end repair enzyme mix (Illumina) and incubated (30 min, 30 °C). End repaired library molecules bound to M280 beads were washed as described (Illumina). A-Tailing and adapter ligation were performed according to manufacturer’s instructions (Illumina). For PCR amplification, the beads were resuspended in a reaction mixture (20 μl nuclease-free water, 25 μl 2x Kappa HiFi (Kappa Biosystems, London, UK), 5 μl Illumina Primer Cocktail) and amplified (98 °C for 3 min, 12 cycles of 98 °C for 10 s, 60 °C for 30 s, 72 °C for 30 s followed by 72 °C for 5 min and storage of the sample at 4 °C). Beads were removed by magnetic separation and 45 μl of the products were transferred to a 2 ml DNA Lobind Eppendorf tube. The DNA was precipitated by addition www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 6 of 31.5 μl CleanPCR beads (GC Biotech, Alphen aan den Rijn, The Netherlands). The beads were washed twice with 100 μl 70% ethanol, and the final library was eluted in 20 μl resuspension buffer (GC biotech). The DNA concentration was determined (Qubit, DNA HS assay, Invitrogen), followed by analysis using the Agilent Bioanalyser (High sensitivity chip, Agilent, Stockport, UK). Up to 12 mate-pair libraries were pooled in an equimolar manner and measured using the Kappa qPCR Illumina quantification kit. Sequencing-by-synthesis of pooled mate-pair libraries was performed using an Illumina HiSeq2500 device (rapid run mode, 2 × 150 cycles paired-end, single indexing reads) according to manufacturer’s instructions (Data Citation 26, Data Citation 27). Sequence assembly of individual BACs Assembly of gene-containing BACs (UCR/JGI). A total of 15,661 gene-bearing BACs were paired-end sequenced (2 × 100 cycles) using the Illumina HiSeq2000 platform (Illumina, Inc., San Diego, CA, USA) applying a combinatorial pooling design 24 , as described in Munoz-Amatriain et al. 13 . Reads were quality trimmed, deconvoluted, and then assembled BAC-by-BAC using Velvet version 1.2.09 (ref. 25) with the parameter k set to 45. Sequences of an additional 50 randomly chosen BACs included in Munoz-Amatriain et al. 13 were derived using the Sanger method by Jane Grimwood (US Department of Energy Joint Genome Institute) and Jeremy Schmutz (HudsonAlpha Institute for Biotechnology), including shatter and transposon sequencing. The assignment of BACs to chromosome arms/pericentromeric regions was performed using CLARK 26 , an accurate k-mer-based classification method that is much faster than BLASTN or MegaBLAST. CLARK makes assignments by using a prebuilt database of k-mers that are specific to each chromosome arm/peri-centromeric region. Assembly of MTP BACs from barley chromosomes 1H, 3H, 4H, 6H and 7H (FLI and IPK). A total of 10,148 BACs mainly originating from barley chromosome 3H were sequenced on the Roche 454 system. Reads were deconvoluted and assigned to individual BACs 16 . Reads were quality trimmed according to the manufacturer’s recommendations. Reads were screened for E. coli and vector sequences with MegaBLAST 27 . Assemblies were then constructed from the clean reads using the MIRA software 28 as described in Steuernagel, et al. 16 and Taudien, et al. 29 . A total of 41,004 BACs were sequenced on Illumina machines (mainly HiSeq2000) in pools of up to 672 individually barcoded BAC clones. Paired-end reads were quality trimmed with the CLC toolkit and screened for E. coli and vector sequences with MegaBLAST. Assemblies were obtained by running CLC Assembly Cell Version 4.0.6 beta with default parameters. Contigs derived with low read coverage as well as contigs smaller than 500 bp were removed using the criteria described in Beier, et al. 17 . The resultant contigs were then compared to NCBI’s nucleotide database using MegaBLAST to check for possible contamination. Contigs with non-plant hits were either completely removed or trimmed. Scaffolding of MTP BACs from barley chromosomes 1H, 3H, 4H, 6Hand7H(FLIandIPK).Scaffolding was performed as described in Beier et al. 17 Briefly, mate-pair reads were mapped against the concatenated assemblies of up to 384 BACs using BWA mem version 0.7.4 (ref. 30) with default parameters. Only read pairs mapping uniquely (minimal mapping quality of Q40) to different contigs of the same BAC assembly were retained. These reads were used to scaffold individual BACs using SSPACE version 3.0 Standard 31 . If multiple mate-pair libraries were present (MiSeq mate-pair reads as well as HiSeq2000 mate-pair reads) an iterative scaffolding procedure 17 was used. Assembly of MTP BACs from barley chromosome 5H (BGI). Obtained raw sequence reads from 5H MTP BACs were filtered to generate high-quality reads by the following criteria: (1) reads containing more than 2% of Ns or with poly-A structures were removed; (2) reads with ≥40% low quality bases for short insert size libraries (60% for large insert size libraries) were excluded; (3) reads containing adapters were removed; (4) PCR duplicates were detected and excluded; (5) removal of reads contaminated by E. coli, vector sequences or phage sequences. High-quality reads were then used for assembly. BACs were assembled using SOAPdenovo version 2.01 (ref. 32) multiple times using different k and m values (main parameter in SOAPdenovo assembly). In total each BAC was assembled 45 times (k from 33 to 66, only odd numbers and m from 1 to 3). The N50 was examined for each assembly and the assembly with the largest N50 was retained as the final assembly result for each BAC. Scaffolding of MTP BACs from barley chromosomes 5H (BGI). Assemblies from paired-end sequences were used as reference for mapping 2, 5 and 10 kb mate-pair reads obtained from barley genomic WGS data with SOAPaligner/soap2 version 2.21 with parameters –p6–v3–R. Mate-pair read pairs mapped in this fashion were used in conjunction with the corresponding paired-end read pairs to re-assemble each BAC using SOAPdenovo version 2.01 as described above. Assembly of MTP BACs from barley chromosomes 2H and ‘0H’(EI). Minimal tiling path BACs from (i) barley chromosomes 2H or from (ii) fingerprinted contigs not assigned to chromosomes (termed ‘0H’) were sequenced. After demultiplexing, sample quality control (QC) information was generated using FastQC 33 . Contamination screening was carried out using Kontaminant 34 . Reads were screened using a k-mer size of 21 against a range of potential contaminants (Phi X, E. coli,Enterobacter cloacae www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 7 genomic DNA and BAC vector) and contaminated reads or reads with quality values o30 were removed. ABySS assembler (v1.5.1) 35 was used to assemble the filtered paired-end reads of each BAC individually (k-71, l-91 b-0). Paired-end contigs were compared to NCBI’s NR database using BLAST to check for hits to non-plant organisms using e-value 1e-4 as threshold. The obtained hits were compared to NCBI taxonomy using ‘fastacmd’to obtain common names used to check for any non-plant hits. Scaffolding of MTP BACs from barley chromosomes 2H and ‘0H’(EI). Illumina Nextera mate-pair libraries were created from pools of 384 BACs. After quality checking the reads using PAP 34 , the reads were merged using FLASH (version 1.2.9) 36 . Nextclip (v0.8) 37 was run on the flashed reads to trim the junction adapters. A k-mer-based approach was used to assign mate-pair reads to individual BACs with KAT (v1.0.4) (https://github.com/TGAC/KAT). Scaffolding and gap closing were performed on each BAC individually using an in-house shell script (available from GitHub: https://github.com/DhSaTGAC/ BAC-assembly-pipeline.git). SOAPdenovo scaffolder version 2.01 (ref. 38) was applied to scaffold the ABySS paired-end contigs using the k-mer classified mate-pair reads with parameters k=41, -G 30, -F, -w and -L 100. The resulting scaffolds were then edited to replace long stretches (>20) of C/G with ‘N’characters as SOAP is known to substitute ‘N’s within paired-end contigs to C/G. The scaffolds were then passed through GapCloser (v1.12-r6), a SOAP2 module, to fill in long stretches of ‘N’s produced during the scaffolding steps. Contigs and scaffolds shorter than 500 bp were removed to produce the final assembly per BAC. Splash contamination checks of MTP BACs from barley chromosomes 2H and ‘0H’(EI). The raw reads within each plate were aligned to one side of the vector sequence adjacent to the restriction enzyme cut site using exonerate 39 . Substrings of size 20 bp were extracted from aligning reads containing the BAC sequence adjacent to the vector sequence. Flanking sequences from each BAC were clustered based on a Hamming distanceo3 and consensus sequences generated to account for sequencing errors. These were compared with neighboring wells to check for potential contamination caused by splash during lab processing steps. Where contamination between neighboring wells was indicated, the assembled contigs from each BAC in question were aligned in a pairwise fashion using exonerate and the total percentage of similar sequence (≥99% identity) was computed. In cases where neighboring BACs shared more than 10% similar sequence, both BACs were resequenced. Pseudomolecule construction Initial contamination removal. Sequence assemblies of 66,586 MTP clones, 5,468 non-MTP BACs and 15,044 gene-bearing clones 13 (total number of unique BACs: 87,098) were combined into a single FASTA file (Data Citation 28,Data Citation 29,Data Citation 30). If a clone had two or more independent sequence assemblies, we selected the one with the largest N50 value for further analyses. BAC assemblies were aligned to a custom library of potential contaminants (Data Citation 31) including phages, bacterial and vector sequences using megablast 27 . Regions aligning to contaminants (criteria: (alignment length ≥500 bp AND identity ≥80%) OR (identity ≥90%)) were removed from the assembly using UNIX scripts and BEDTools 40 . Sequences shorter than 500 bp or consisting of less than 500 proper nucleotides (ACGT characters) after contamination removal were discarded. This step removed 55.5 Mb (0.5%) of the assembled BAC sequence. Sequence alignment of BACs sequences and overlap detection. After contamination removal, a set of 87,075 BAC assemblies (Table 1, Data Citation 32) was aligned against itself using megablast 27 with a word size of 44, retaining only alignments with identity ≥99% and alignment length ≥500 bp. Two sets of overlaps (stringent and permissive) between BACs were defined from the BLAST results of all BACs against each other. Pairs of BACs were considered as potentially overlapping under stringent criteria if there was at least one high-scoring pair (HSP) with alignment length ≥5 kb and identity ≥99.8%. Under permissive criteria, we required at least one HSP with alignment length ≥2 kb and identity ≥99.5%. For all pairs of potentially overlapping BACs (under either set of criteria), the size of their overlapping regions was determined using UNIX scripts and BEDTools 40 as the extent of non-redundant regions in the BAC sequences (i.e., contigs or scaffolds) contained in HSPs ≥500 bp and identity ≥99.5% between BAC sequences having at least one HSP with alignment length ≥5 kb and identity ≥99.8% (stringent criteria) or alignment length ≥2 kb and identity ≥99.5% (permissive criteria). HSPs less than 200 bp apart were combined into one with BEDTools (command ‘merge’). BAC overlap information was imported into the R statistical environment 41 for use in genetic anchoring and merging sequence assemblies of adjacent BAC clones (see section ‘Construction of the BAC overlap graph’). Alignment of BACs to the BioNano map of barley cv. Morex. An optical map of the genome of barley cv. Morex was generated using the Irys platform of BioNano Genomics using Nt.BspQI as the nicking enzyme. Further details of the optical map procedure are described in Mascher et al. 42 An in silico BspQI digest was performed with the Knickers software (http://www.bionanogenomics.com) using default parameters. Restriction maps of BAC sequences were aligned to the BioNano map of barley cv. Morex 42 (Data Citation 33) with IrysView software 43 (http://www.bionanogenomics.com) using the www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 8 command line tool RefAligner (version 3827) with the following parameters ‘-M 2 -T 1e-4 -extend 1 -biaswt 0’to report all alignments with a confidence score ≥4. Construction of the updated POPSEQ map of the Morex x Barke mapping population. An ultradense linkage map had been constructed previously 5 by shallow whole-genome shotgun sequencing of 90 recombinant inbred lines (RILs) derived from a cross between the barley cultivars Morex and Barke. We wished to increase the resolution of this map by reducing the average fraction of missing data per SNP marker. Towards this aim, we sequenced the existing Illumina paired-end libraries of 87 RILs to higher coverage (2–3x) and combined them (Data Citation 34) with the existing read data set 5 (ENA accession: ERP002184). Map construction followed the procedures described in Chapman et al. 44 . Reads were aligned to the whole-genome shotgun assembly of barley cv. Morex 4 (NCBI accession: CAJW01) with BWA mem version 0.7.5a (ref. 45). Sorting, conversion to BAM format and removal of duplicate reads was done with PicardTools version 1.100 (http://broadinstitute.github.io/picard/). Variant detection and genotype calling were performed with SAMTools version 0.1.19 (commands ‘samtools mpileup –BD’and ‘bcftools view –cvg’). The resultant VCF file was filtered using an AWK script (Supplementary Text S3 of Mascher et al. 2013 (ref. 46)). Homozygous genotype calls were set to missing if their read depth was 0 or their genotype quality below 3. Heterozygous genotype calls were set to missing if their read depth was below 3 or their genotype quality below 5. Variants with (i) a quality scores below 40, (ii) more than 10% heterozygous genotype calls, (iii) more than 90% missing data after genotype call filtering, or (iv) a minor allele frequency below 5% were discarded. SNP information was aggregated at the contig level to derive consensus genotypes as described in the section ‘Framework map construction’in the Methods section of Chapman et al. 44 For map construction with MSTMap 47 , the population type ‘RIL8’was used. Additional contigs were inserted into the framework map as described in Chapman et al. 44 (section ‘Anchoring scaffolds onto the framework map’) using previously published read data 5 . Variant calling and map construction were done for the Oregon Wolfe Barley (OWB) doubled haploid population using the same procedures with the following two changes: (i) heterozygous genotype calls were excluded and (ii) the population type ‘DH’was used for map construction with MSTMap 47 . Map positions in the OWB map were interpolated into the Morex x Barke map using loess regression in R 41 . A consensus position was derived as follows: if map positions disagreed by more than 5 cM in both maps, a contig was considered unanchored; otherwise, the Morex x Barke position was preferred if available. The final map assigned genetic positions to 791,176 WGS contigs (Table 2, Data Citation 35), compared to 723,499 anchored contigs in the original POPSEQ map 5 . Genetic anchoring of single BAC clones. The genetic positions of Morex WGS contigs in the updated POPSEQ map were lifted to BAC sequences via sequence alignment. The set of all contigs of the wholegenome shotgun assembly of barley cv. Morex 4 (NCBI accession: CAJW01) was aligned to all BAC assemblies with megablast 27 using a word size of 44 and retaining only alignments with identity ≥99.8% and alignment length ≥1,000 bp. For each BAC clone, the genetic positions of WGS contigs aligning to its constituent sequences were tabulated and a genetic position of a clone was derived using a majority rule with functions of the R package ‘data.table’(https://cran.r-project.org/web/packages/data.table/index. html). Ninety per cent of contigs assigned to a BAC had to originate to the major chromosome and the standard deviation of genetic positions had to be ≤3 cM. BACs without alignments to anchored WGS contigs were considered as unanchored; those not meeting the consistency criteria were flagged as ‘inconsistently anchored’. In the second step, unanchored clones were positioned by utilizing positional information from neighboring BACs. We considered as neighbors of a given clone B all those BACs that overlapped for at least 10% of their assembled lengths with clone B. The genetic position of an MTP chromosome no of. BACs in MTP no. of sequenced BACs no. of anchored BACs*average no. of sequences average N50 (kb) 1H 6,993 6,983 (99.9%) 6,410 (91.8%) 7.6 81.2 2H 9,061 8,969 (99.0%) 8,195 (91.4%) 9.9 104.5 3H 8,841 8,807 (99.6%) 8,303 (94.3%) 7.7 87.5 4H 8,314 8,306 (99.9%) 7,783 (93.7%) 6.7 91.2 5H 8,426 8,358 (99.2%) 7,573 (90.6%) 9.7 72.2 6H 8,305 7,886 (95.0%) 6,476 (82.1%) 7.4 70.7 7H 8,576 7,970 (92.9%) 6,842 (85.8%) 8.5 65.5 ‘0H’ † 8,256 8,031 (97.3%) 6,714 (83.6%) 7.6 83.6 Non-MTP —21,765 20,397 (93.7%) 14.5 33.7 Total 66,772 87,075 78,693 (90.4%) 9.8 70.3 Table 1. BAC assembly and anchoring statistics. *Number and percentage of BAC clones that have been assigned genetic positions in the POPSEQ map. † BAC clones in physical contigs that had not been assigned to chromosomes. www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 9 length ≥500 bp and an alignment identity ≥99.5%. Regions of the old non-redundant sequence covered by C (as determined by commands of BEDTools 40 suite) were removed and contig C was added instead. This procedure was performed for each C2-type contig. Next, we queried the GMAP alignments for genes that had no alignments to the non-redundant sequence, but were represented either in (a) the Morex WGS contigs or in (b) sequences of BACs excluded from the overlap analysis. We considered sequence of type (a) and (b) as ‘additional gene-bearing sequences’. We aligned these additional gene-bearing sequences to the non-redundant sequence with megablast 27 using a word size of 44 and considering only high-scoring pairs with an alignment length ≥500 bp and an alignment identity ≥99.5%. Regions covered by the non-redundant sequence under these alignment criteria were subtracted from the additional gene-bearing sequences and sequence fragments with a length ≥500 bp were added to the non-redundant sequence. Final contamination removal. We identified regions in the non-redundant sequence that were not covered by whole-genome shotgun reads of cv. Morex. Alignment of WGS reads and read depth calculation were done as described in the section ‘Alignment of Hi-C data to restriction fragments’. Regions of the non-redundant sequence not covered by Morex WGS reads and with a length ≥500 bp were extracted using UNIX command line tools and BEDTools 40 (command ‘getfasta’). The extracted sequences were aligned to the NCBI NT database with megablast 27 using a word size of 44 and requiring the high-scoring pairs to have a length of at least 100 bp and an alignment identity ≥80%. We retained only hits whose description in the NCBI NT database did not match the following regular expression (R syntax) representing a list of common and taxonomic names of plant species: ‘Hordeum|Triti|Populus|Aegilops|Avena|Alnus|A\\.squarrosa|Morus|Nelumbo|Brassica|Cucumis|Citrus| Camelina|Fragaria|Lotus|Tarenaya|Spartina|Eucommia|Sorghum|Corylus|Theobroma|Phaseolus|Barley| Trifolium|Elymus|Brachypodium|Beta vulgaris|Ricinus|Licania|Phoenix|H\\.vulgare|Pyrus|Malus|Prunus| Saccharum|Hypericum|Wheat|Oryza|hloroplast|Secale|Vitis|Quercus’ Regions overlapping the BLAST hits passing these filters were cut from the non-redundant sequence with BEDTools 40 (command ‘subtract’). Sequences shorter than 500 bp after the removal of contaminant sequences were discarded. This step removed 5 Mb (0.1%) of the assembled sequence. Construction of pseudomolecule sequences for chromosome 1H—7H and chrUn. We constructed pseudomolecules of the seven barley chromosomes by placing the sequence fragments of single BAC assemblies that constitute the non-redundant sequence according to the Hi-C map positions of the BAC overlap clusters these fragments belong to. Sequences not anchored by Hi-C were placed on chrUn (‘chromosome unassigned’). The order of clusters was taken from the Hi-C map. BACs within the same cluster were ordered according to the minimum spanning tree of the BAC overlap graph of the cluster and oriented relative to the telomeres using the Hi-C orientation of the cluster if available. The relative order of sequence fragments originating from the same BAC bin (see section ‘Construction of the BAC overlap graph’) could not be determined so that the placement of sequences within a BAC bin (average size: 70 kb) is arbitrary. ChrUn is composed of (i) sequence fragments originating from BAC overlap clusters not placed in the Hi-C map, or (ii) gene-bearing fragments of BAC sequences and Morex WGS contigs selected in addition to the non-redundant sequence (see section Identification of additional genebearing sequences). A gap of 100 N characters was inserted between adjacent sequence fragments. Pseudomolecules of all chromosomes and chrUn were combined into a single FASTA file (Data Citation 42). To accommodate limitations of the Sequence/Alignment Map format (see Usage Notes) split pseudomolecules with a size below 512 Mb were constructed by breaking pseudomolecules arbitrarily at breaks between sequence contigs (Data Citation 43, Data Citation 44). A BED file indicating the placement of BAC sequence fragments, Morex WGS contigs and intercalating gaps in the (split) pseudomolecules is available for download (Data Citation 45, Data Citation 46). A tabular summary of the positional information incorporated into pseudomolecules is given in Data Citation 41. Masking of residual redundancy Residual redundancy arising from undetected overlaps between adjacent BACs was detected and masked by aligning the pseudomolecules sequence to itself with megablast 27 . Genomic intervals contained in BLAST hits with a length ≥5 kb and an identity ≥99.8% were considered as potentially redundant (PR) regions. PR regions were classified to decide which sequence of a redundant pair to mask: (i) PR regions assigned to chromosomal pseudomolecules (as opposed to chrUn), but having BLAST hits only to other chromosomes were considered as originating from chimeric BAC assemblies incorporating unrelated sequences from different chromosomes and masked with Ns; (ii) an analogous procedures was used to find intrachromosomal chimeras based on Hi-C map information; (iii) PR regions on chrUn that had alignments to regions on chromosomal pseudomolecules were masked, (iv) for other PR regions one sequence of a redundant pair was chosen arbitrarily. Positions of masked regions on the (split) pseudomolecules were written into a BED file (Data Citation 47, Data Citation 48). Masking was done with BEDTools 40 (command ‘mask’) overwriting nucleotides in redundant intervals with N characters. Masked versions of the (split) pseudomolecules are provided as Data Citation 49, Data Citation 50). www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 16 POPSEQ genetic map based on pseudomolecule sequence After the construction of the map-based reference sequence, we constructed an updated high-resolution genetic map of the Morex x Barke population to validate the order of genetic map in the reference Figure 2. Collinearity between the Hi-C map and two genetic maps. The positions of genetic markers (x-axis) are plotted against their genetic positions (y-axis) in a GBS map (top row) and a POPSEQ map (bottom row) of the Morex x Barke recombinant inbred lines. www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 17 sequence. Raw reads (see section ‘Construction of the updated POPSEQ map of the Morex x Barke mapping population’) were aligned to the barley pseudomolecules with BWA mem (version 0.7.12) 45 . Checking mated mapped paired reads, sorting, conversion to BAM format and marking of duplicate read pairs were done with PicardTools version 2.300 (http://broadinstitute.github.io/picard/). Variant detection and genotype calling were performed using GATK Toolkit version 3.3.0 (command ‘HaplotypeCaller’) 57 . A total of five RILs with >3% heterozygous variants were removed. A variant position was removed if more than 10% of all samples were called heterozygous, there were more than 80% missing data, or the minor allele frequency (in the non-missing data) was smaller than 5%. SNP information was aggregated at the contig level to derive consensus genotype blocks with false discovery rate calculated based on the quality of each variant call in the block. High-confidence genotype blocks were obtained based on a Bonferroni correction threshold. Given the fact that the length of crossover tracts is significantly larger than that of non-crossover tracts and non-crossover tracts would enlarge the genetic distance artificially, we only retained high-confidence genotype blocks with more than 1 Mb tract length, which are likely to be derived from crossovers. Representative non-redundant genomic variants of high-confidence genotype blocks were extracted and used for the construction of a high-resolution map through MSTMap 47 . We further anchored all remaining markers to the genetic map by the C program ‘canchor’ 5 . The final POPSEQ map consisted of 9,012,742 SNP variants defined on the pseudomolecule sequence Data citation 51). Representation of full-length cDNAs The representation of gene models in the whole-genome genome assembly of barley cv. Morex 4 and in the pseudomolecules was compared by aligning a set of 22,651 publicly available full-length cDNAs 55 to the assemblies using the GMAP splice aligner software 56 . The GMAP alignment output was then filtered. If a full-length cDNA had multiple hits, only the hit with the highest % identity was considered. Hits were further filtered by identity (≥98%) and coverage ( ≥95%). This resulted in a set of hits representing genes recovered intact on a single genomic contig/chromosome. Code availability R and shell source code for the construction of the BAC overlap graph and the Hi-C map is provided as Data Citation 52. Code can be re-used under the terms of the MIT license. Data Records BAC sequence raw data was submitted to the European Nucleotide Archive (ENA) (Data Citation 1, Data Citation 2, Data Citation 3, Data Citation 4, Data Citation 5, Data Citation 6, Data Citation 7, Data Citation 8, Data Citation 9, Data Citation 10, Data Citation 11, Data Citation 12, Data Citation 13, Data Citation 14, Data Citation 15, Data Citation 16, Data Citation 17, Data Citation 18, Data Citation 19, Data Citation 20, Data Citation 21, Data Citation 22, Data Citation 23, Data Citation 24, Data Citation 25, Data Citation 26, Data Citation 27). BAC assemblies were submitted to ENA or NCBI (Data Citation 28, Data Citation 29). Raw data for POPSEQ (Data Citation 35), GBS (Data Citation 38) and Hi-C mapping (Data Citation 40) were submitted to ENA. Processed datasets are accessible as Figure 3. Collinearity between the Hi-C map and a cytogenetic map of chromosome 3H. Dots mark the positions of probes in the cytogenetic map (x-axis) and the Hi-C-derived pseudomolecule (y-axis). A linear regression line (red) was fitted with the R function lm(). Note that cytogenetic data is not available for distal regions because probes were designed only for non-recombining peri-centromeric regions 61 . www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 18 QRLPDAAGSSAEEHSGQDKLLIVVTPTR ARASQAYYLSRMGQTLRLVRPPVLWVVV EAGKPTPEAALELRRTAVMHRYVGCCDA LNASASPAVDFRPHQLNAGLEVVENHRL DGVVYFADEEGVYSLPLFDRLRQIRRFG TWPVPTISDGGHGVVLEGPVCKQNQVVG WHTSGDANKLQRFHVAMSGFAFNSTMLW DPRLRSHKAWNSIRHPEMVEQGFQGTTF VEQLVEDESQMEGIPADCSQIMNWHVPF GSESPVYPKGWRSAANLDVIIPLK Figure 4. Accessing sequence and positional information with the barley genome explorer (BARLEX). The barley pseudomolecule data was imported into BARLEX, where it is directly linked to the IPK Barley BLAST server. Users can paste a nucleotide or amino acid sequence (1) into the BARLEX input query form and select reference database such as pseudomolecules sequence, the set of all BAC assemblies or annotated genes (2). The sequence is then transferred to the IPK barley BLAST Server (3). The web page with the BLAST results (4) contains references to BARLEX information pages for different structural units (BAC sequence contigs, BAC, BAC cluster, chromosomal Hi-C map). For example, the pages of BAC sequence contigs visualize the repeat content based on genome-wide k-mer histograms (5) and are linked to a graph-based visualization (6) of the entire BAC assembly. Summary statistics and positional information of BAC clusters are presented in tables that can be searched, sorted and subsetted using user-defined criteria (7). Users can convert pseudomolecule coordinates (AGP positions) to intervals in the underlying BAC sequence assemblies (8). www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 19 Digital Object Identifiers (DOIs) in the Plant Genomics and Phenomics Research Data Repository 58 (Data Citation 30, Data Citation 31, Data Citation 32, Data Citation 33, Data Citation 34, Data Citation 36, Data Citation 37, Data Citation 39, Data Citation 41, Data Citation 42, Data Citation 43, Data Citation 44, Data Citation 45, Data Citation 46, Data Citation 47, Data Citation 48, Data Citation 49, Data Citation 50, Data citation 51, Data Citation 52). DOIs were registered with e!DAL 59 . Technical Validation Collinearity between genetic maps and pseudomolecules To validate the order of scaffolds in the Hi-C map, we compared the order of genetic marker loci in the Hi-C-derived pseudomolecules to their positions in linkage maps. First, we used genotyping-bysequencing (GBS) 11,50 to type single-nucleotide polymorphisms (SNPs) segregating in a bi-parental population comprising 2,398 recombinant inbred lines (RILs). A total of 2,637 SNPs were detected by aligning GBS reads and calling variants and genotypes using a previously published pipeline 46 . Second, we reanalysed WGS re-sequencing data of a subset of the same population (POPSEQ data) comprising 90 RILs. Construction of a framework linkage map and insertion of additional markers were performed essentially as described by Chapman et al. 44 A dot plot comparison of physical and genetic SNP positions revealed that marker orders were highly collinear between the pseudomolecules and both the GBS and POPSEQ map of the Morex x Barke population (Fig. 2). Collinearity between a cytogenetic map and the pseudomolecule of chromosome 3H We could not validate the order of BAC overlap clusters in the large peri-centromeric regions because of severely repressed recombination 3,60 . Therefore, we compared the order of probes mapped by fluorescence in-situ hybridization to chromosomal locations on chromosome 3H and their corresponding sequences in the pseudomolecule of 3H. Since probes were derived from BAC sequences associated with physical contigs, their position from the reference sequence could be determined from the BAC overlap graph. The comparison showed that the cytogenetic and Hi-C maps were highly collinear in pericentromeric regions of chromosome 3H (Fig. 3). Representation of full-length cDNAs To assess the completeness of our assembly, we checked for the presence of high-confidence transcript sequences. The representation of gene models in the whole-genome shotgun assembly of barley cv. Morex 4 and in the map-based reference assembly was compared by aligning a set of 22,651 publicly available full-length cDNAs 55 of barley cv. ‘Haruna Nijo’. After aligning and filtering, 18,062 (79.74%) intact full-length cDNAs were found in the pseudomolecules, whereas only 10,496 (46.33%) were recovered in the whole-genome assembly. This increase in the number of correctly represented full-length cDNAs vindicates the effort invested in the map-based assembly. Nevertheless, a significant proportion of genes remain fragmented even in the pseudomolecule assembly (20.26%), and presumably these largely represent difficult to assemble genes that contain e.g., microsatellites, long homopolymer stretches and other difficult features, and/or form part of complex gene families that are difficult to resolve. It is likely that only longer read technologies such as Pacific Biosciences (http://www.pacb.com) or Oxford Nanopore (https://www.nanoporetech.com) will be able to resolve these more difficult cases. Further results on gene space completeness based on an automated gene annotation of the pseudomolecules, and on the representation of repetitive elements are described elsewhere 42 . Usage Notes Positional information for BAC sequences, physical contigs and WGS contigs can be accessed via the barley genome explorer BARLEX (Fig. 4). BLAST searches against the barley pseudomolecules can also be carried out in BARLEX. We note that processing BAM files with short read alignments to the full pseudomolecules with commonly used tools such as SAMtools 52 or BEDTools 40 may not work as expected because of restrictions on the chromosome size (512 Mb) for indexing file in Sequence Alignment/Map (SAM) format 52 . To circumvent this issue, we have split the pseudomolecules into two part and provide (i) a FASTA file with split pseudomolecules (Data Citation 44) along the with the intact sequences and (ii) a BEDfile to convert between full and split pseudomolecule coordinate (Data Citation 43) Alternatively, the CRAM format (https://samtools.github.io/hts-specs/CRAMv3.pdf) may be used instead of the BAM format. We note that the orientation of sequence contigs within individual BACs in the pseudomolecules is arbitrary, thus the order and orientation of sequences in the pseudomolecules is accurate only up to resolution of ~100 kb. References 1. Schulte, D. et al. The international barley sequencing consortium--at the threshold of efficient access to the barley genome. Plant physiology 149, 142–147 (2009). 2. Schulte, D. et al. BAC library resources for map-based cloning and physical map construction in barley (Hordeum vulgare L). BMC genomics 12, 247 (2011). 3. Ariyadasa, R. et al. A sequence-ready physical map of barley anchored genetically by two million single-nucleotide polymorphisms. Plant physiology 164, 412–423 (2014). 4. International Barley Genome Sequencing Consortium. A physical, genetic and functional sequence assembly of the barley genome. Nature 491, 711–716 (2012). www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 20 5. Mascher, M. et al. Anchoring and ordering NGS contig assemblies by population sequencing (POPSEQ). The Plant Journal 76, 718–727 (2013). 6. Lander, E. S. et al. Initial sequencing and analysis of the human genome. Nature 409, 860–921 (2001). 7. Schnable, P. S. et al. The B73 maize genome: complexity, diversity, and dynamics. Science 326, 1112–1115 (2009). 8. Lam, E. T. et al. Genome mapping on nanochannel arrays for structural variation analysis and sequence assembly. Nature biotechnology 30, 771–776 (2012). 9. Lieberman-Aiden, E. et al. Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science 326, 289–293 (2009). 10. Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nature biotechnology 31, 1119–1125 (2013). 11. Poland, J. A., Brown, P. J., Sorrells, M. E. & Jannink, J.-L. Development of high-density genetic maps for barley and wheat using a novel two-enzyme genotyping-by-sequencing approach. PLoS ONE 7, e32253 (2012). 12. Colmsee, C. et al. BARLEX—the Barley Draft Genome Explorer. Mol Plant 8, 964–966 (2015). 13. Munoz-Amatriain, M. et al. Sequencing of 15 622 gene-bearing BACs clarifies the gene-dense regions of the barley genome. Plant Journal 84, 216–227 (2015). 14. Pasquariello, M. et al. The barley Frost resistance-H2 locus. Functional & integrative genomics 14, 85–100 (2014). 15. Meyer, M., Stenzel, U. & Hofreiter, M. Parallel tagged sequencing on the 454 platform. Nature protocols 3, 267–278 (2008). 16. Steuernagel, B. et al. De novo 454 sequencing of barcoded BAC pools for comprehensive gene survey and genome analysis in the complex genome of barley. BMC genomics 10, 547 (2009). 17. Beier, S. et al. Multiplex sequencing of bacterial artificial chromosomes for assembling complex plant genomes. Plant biotechnology journal 14, 1511–1522 (2016). 18. Sambrook, J. & Russell, D. W. Molecular cloning: a laboratory manual. 3rd edition (Coldspring-Harbour Laboratory Press, 2001). 19. Aird, D. et al. Analyzing and minimizing PCR amplification bias in Illumina sequencing libraries. Genome biology 12, R18 (2011). 20. Quail, M. A. et al. A large genome center’s improvements to the Illumina sequencing system. Nature methods 5, 1005–1010 (2008). 21. Asan et al. Paired-end sequencing of long-range DNA fragments for de novo assembly of large, complex Mammalian genomes by direct intra-molecule ligation. PLoS ONE 7, e46211 (2012). 22. Meyer, M. & Kircher, M. Illumina sequencing library preparation for highly multiplexed target capture and sequencing. Cold Spring Harb Protoc 2010, pdb prot5448 (2010). 23. Adey, A. et al. Rapid, low-input, low-bias construction of shotgun fragment libraries by high-density in vitro transposition. Genome biology 11, R119 (2010). 24. Lonardi, S. et al. Combinatorial pooling enables selective sequencing of the barley gene space. PLoS computational biology 9, e1003010 (2013). 25. Zerbino, D. R. & Birney, E. Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Genome research 18, 821–829 (2008). 26. Ounit, R., Wanamaker, S., Close, T. J. & Lonardi, S. CLARK: fast and accurate classification of metagenomic and genomic sequences using discriminative k-mers. BMC genomics 16, 236 (2015). 27. Zhang, Z., Schwartz, S., Wagner, L. & Miller, W. A greedy algorithm for aligning DNA sequences. Journal of computational biology: a journal of computational molecular cell biology 7, 203–214 (2000). 28. Chevreux, B., Wetter, T. & Suhai, S. in German conference on bioinformatics (1999); 45–56. 29. Taudien, S. et al. Sequencing of BAC pools by different next generation sequencing platforms and strategies. BMC research notes 4, 411 (2011). 30. Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics 25, 1754–1760 (2009). 31. Boetzer, M., Henkel, C. V., Jansen, H. J., Butler, D. & Pirovano, W. Scaffolding pre-assembled contigs using SSPACE. Bioinformatics 27, 578–579 (2011). 32. Brenchley, R. et al. Analysis of the bread wheat genome using whole-genome shotgun sequencing. Nature 491, 705–710 (2012). 33. Andrews, S. FastQC: a quality control tool for high throughput sequence data. Available online at: http://www.bioinformatics. babraham.ac.uk/projects/fastqc (2010). 34. Leggett, R. M., Ramirez-Gonzalez, R. H., Clavijo, B. J., Waite, D. & Davey, R. P. Sequencing quality assessment tools to enable data-driven informatics for high throughput genomics. Frontiers in genetics 4, 288 (2013). 35. Simpson, J. T. et al. ABySS: a parallel assembler for short read sequence data. Genome research 19, 1117–1123 (2009). 36. Magoc, T. & Salzberg, S. L. FLASH: fast length adjustment of short reads to improve genome assemblies. Bioinformatics 27, 2957–2963 (2011). 37. Leggett, R. M., Clavijo, B. J., Clissold, L., Clark, M. D. & Caccamo, M. NextClip: an analysis and read preparation tool for Nextera Long Mate Pair libraries. Bioinformatics 30, 566–568 (2014). 38. Luo, R. et al. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. GigaScience 1, 18 (2012). 39. Slater, G. S. & Birney, E. Automated generation of heuristics for biological sequence comparison. BMC bioinformatics 6, 31 (2005). 40. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841–842 (2010). 41. R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing, 2015). 42. Mascher, M. et al. A chromosome conformation capture ordered sequence of the barley genome. Nature doi:10.1038/nature22043 (2017). 43. Cao, H. et al. Rapid detection of structural variation in a human genome using nanochannel-based genome mapping technology. GigaScience 3, 1 (2014). 44. Chapman, J. A. et al. A whole-genome shotgun approach for assembling and anchoring the hexaploid bread wheat genome. Genome biology 16, 26 (2015). 45. Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. Preprint at https://arxiv.org/pdf/ 1303.3997v2.pdf (2013). 46. Mascher, M., Wu, S., Amand, P. S., Stein, N. & Poland, J. Application of genotyping-by-sequencing on semiconductor sequencing platforms: a comparison of genetic and reference-based marker ordering in barley. PLoS ONE 8, e76925 (2013). 47. Wu, Y., Bhat, P. R., Close, T. J. & Lonardi, S. Efficient and accurate construction of genetic linkage maps from the minimum spanning tree of a graph. PLoS genetics 4, e1000212 (2008). 48. Csardi, G. & Nepusz, T. The igraph software package for complex network research, InterJournal, Complex Systems 1695 (2006). 49. Prim, R. C. Shortest connection networks and some generalizations. Bell system technical journal 36, 1389–1401 (1957). 50. Wendler, N. et al. Unlocking the secondary gene-pool of barley with next-generation sequencing. Plant biotechnology journal 12, 1122–1131 (2014). www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 21 51. Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet. journal 17, 10–12 (2011). 52. Li, H. et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078–2079 (2009). 53. Garrison, E. & Marth, G. Haplotype-based variant detection from short-read sequencing. Preprint at https://arxiv.org/pdf/ 1207.3907v2.pdf (2012). 54. Kalhor, R., Tjong, H., Jayathilaka, N., Alber, F. & Chen, L. Genome architectures revealed by tethered chromosome conformation capture and population-based modeling. Nature biotechnology 30, 90–98 (2012). 55. Matsumoto, T. et al. Comprehensive sequence analysis of 24,783 barley full-length cDNAs derived from 12 clone libraries. Plant physiology 156, 20–28 (2011). 56. Wu, T. D. & Watanabe, C. K. GMAP: a genomic mapping and alignment program for mRNA and EST sequences. Bioinformatics 21, 1859–1875 (2005). 57. DePristo, M. A. et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature genetics 43, 491–498 (2011). 58. Arend, D. et al. PGP repository: a plant phenomics and genomics data publication infrastructure. Database 2016, baw033 (2016). 59. Arend, D. et al. e!DAL--a framework to store, share and publish research data. BMC bioinformatics 15, 214 (2014). 60. Künzel, G., Korzun, L. & Meister, A. Cytologically integrated physical restriction fragment length polymorphism maps for the barley genome based on translocation breakpoints. Genetics 154, 397–412 (2000). 61. Aliyeva-Schnorr, L. et al. Cytogenetic mapping with centromeric bacterial artificial chromosomes contigs shows that this recombination-poor region comprises more than half of barley chromosome 3H. The Plant Journal 84, 385–394 (2015). Data Citations 1. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9062 (2016). 2. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9097 (2016). 3. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9098 (2016). 4. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9099 (2016). 5. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9100 (2016). 6. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9101 (2016). 7. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9102 (2016). 8. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9103 (2016). 9. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9104 (2016). 10. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB8576 (2016). 11. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB8577 (2016). 12. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB8578 (2016). 13. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9619 (2016). 14. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB8579 (2016). 15. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB8580 (2016). 16. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9429 (2016). 17. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9430 (2016). 18. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9431 (2016). 19. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB10963 (2016). 20. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB11489 (2016). 21. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB12096 (2016). 22. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB11758 (2016). 23. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9428 (2016). 24. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB11991 (2016). 25. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB9427 (2016). 26. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB11798 (2016). 27. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB11992 (2016). 28. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB13020 (2016). 29. Muñoz-Amatriaín, M. et al. NCBI BioProject PRJNA198204 (2015). 30. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/21 (2016). 31. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/28 (2016). 32. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/12 (2016). 33. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/31 (2016). 34. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB13028 (2016). 35. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/33 (2016). 36. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/22 (2016). 37. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/30 (2016). 38. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB14130 (2016). 39. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/29 (2016). 40. International Barley Genome Sequencing Consortium. European Nucleotide Archive PRJEB14169 (2016). 41. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/20 (2016). 42. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/34 (2016). 43. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/27 (2016). 44. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/36 (2016). 45. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/23 (2016). 46. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/24 (2016). 47. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/25 (2016). 48. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/26 (2016). 49. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/35 (2016). 50. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/37 (2016). 51. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/17 (2016). 52. International Barley Genome Sequencing Consortium. IPK Gatersleben http://dx.doi.org/10.5447/IPK/2016/19 (2016). Acknowledgements This work was carried out under the auspices of the International Barley Genome Sequencing Consortium and supported from the following funding sources: German Ministry of Education and Research (BMBF) grant 0314000 ‘BARLEX’and 0315954 ‘TRITEX’to M.P., U.S. and N.S and 031A536 www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 22 ‘de.NBI’to U.S. Leibniz Association grant (‘Pakt f. Forschung und Innovation’)‘sequencing barley chromosome 3H’to N.S. and U.S.; Scottish Government/UK Biotechnology and Biological Sciences Research Council (BBSRC) grant BB/100663X/1 to R.W, P.E.H., J.R.; BBSRC grants BB/I008357/1 to M. D.C., M.C. and BB/I008071/1 to P.K.; of Finland grant 266430 and a BioNano grant to A.H.S.; Carlsberg Foundation grant nr. 2012_01_0461 to the Carlsberg Research Laboratory; Grain Research and Development Corporation (GRDC) grant DAW00233 to C.L. and P.L.; Department of Agricultural and Food, Government of Western Australia grant 681 to C.L.; National Natural Science Foundation of China (NSFC) grant 31129005 to C.L. and G.Zhang; NSFC grant 31330055 to G.Zhang.; Czech Ministry of Education, Youth and Sports grant LO1204 to J.D.; National Science Foundation grant DBI 0321756 ‘Coupling EST and Bacterial Artificial Chromosome Resources to Access the Barley Genome’to T.J.C. and S.L.; United States Department of Agriculture (USDA), Agriculture and Food Research Initiative Plant Genome, Genetics and Breeding Program of USDA-CSREES-NIFA grant 2009-65300-05645 ‘Advancing the Barley Genome’and 2011-68002-30029 ‘TriticeaeCAP’to T.J.C., S.L. and G.J.M.; United States National Science Foundation (NSF)-ABI grant DBI-1062301 to T.J.C. and S.L.; University of California grant CA-R-BPS-5306-H to T.J.C and S.L.;National Science Foundation grant DBI 0321756 ‘Algorithms for Genome Assembly of Ultra-deep Sequencing Data’to S.L. Next-generation sequencing and library construction was delivered via the BBSRC National Capability in Genomics (BB/J010375/1) at Earlham Institute (formerly The Genome Analysis Centre) by members of the Platforms and Pipelines group and BBSRC Institute Strategic Programme funding for Bioinformatics (BB/J004669/1) to M.D.C., S.A. and M.C. We gratefully acknowledge: (1) the excellent technical assistance by Susanne König, Manuela Knauft, Uli Beier, Anne Kusserow, Katrin Trnka, Ines Walde, Sandra Driesslein, Cynthia Voss; (2) Doreen Stengel, Anne Fiebig, Thomas Münch, Danuta Schüler and Daniel Arend and Matthias Lange for sequence raw data management and data submission to EMBL/ENA and registration of DOIs; (3) Dr Hélène Berges, Arnaud Bellec and Sonia Vautrin (CNRGV) for management and distribution of barley BAC libraries; (4) Andreas Graner and David Marshall for scientific discussions. Author Contributions BAC sequencing and assembly (1H, 3H, 4H): S.B., A.Himmelbach, S.T., M.F., M.G., M.M., U.S. (co-leader), M.P. (co-leader), N.S. (leader); BAC sequencing and assembly (2H, unassigned): D.S., D.H., S. A. (co-leader), M.D.C. (co-leader), M.C. (co-leader), R.W. (leader); BAC sequencing and assembly (5H, 7H): X.Z., R.A.B., Q.Z., C.T., J.K.M., B.C., G.Zhou, F.D., Y.H., S.Y., S.Cao, S.Wang, X.L., M.I.B., P.L., G.Zhang (co-leader), C.Li (leader); BAC sequencing and assembly (6H): S.B., S.Wang, C.Lin, H.L., U.S., M. H. (co-leader), I.B. (leader); BAC sequencing (gene-bearing): M.M.-A., R.O., S.Wanamaker, S.L. (co-leader), T.J.C. (leader); Optical mapping: A.Hastie, H.Š., J.T., H.S., J.V., S.Chan, M.M., N.S., J.D., A.H.S. (leader); Chromosome conformation capture: A.Himmelbach, S.G., M.M. (co-leader), N.S. (leader); Pseudomolecule construction: M.M. (leader), S.B., C.C., D.B., T.S., P.K., N.S., U.S. (co-leader); Validation: L.L., M.B., L.A.-S., A.Houben, J.A.P., N.S., G.J.M., M.M. (leader). All authors read and commented on the manuscript. Additional information Competing financial interests: The authors declare no competing financial interests. How to cite this article: Beier, S. et al. Construction of a map-based reference genome sequence for barley, Hordeum vulgare L. Sci. Data 4:170044 doi: 10.1038/sdata.2017.44 (2017). Publisher’s note: Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. This work is licensed under a Creative Commons Attribution 4.0 International License. The images or other third party material in this article are included in the article’s Creative Commons license, unless indicated otherwise in the credit line; if the material is not included under the Creative Commons license, users will need to obtain permission from the license holder to reproduce the material. To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0 Metadata associated with this Data Descriptor is available at http://www.nature.com/sdata/ and is released under the CC0 waiver to maximize reuse. © The Author(s) 2017 www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 23 Sebastian Beier 1,* , Axel Himmelbach 1,* , Christian Colmsee 1 , Xiao-Qi Zhang 2 , Roberto A. Barrero 3 , Qisen Zhang 4 , Lin Li 5 , Micha Bayer 6 , Daniel Bolser 7 , Stefan Taudien 8 , Marco Groth 8 , Marius Felder 8 , Alex Hastie 9 , Hana Šimková 10 , Helena Staňková 10 , Jan Vrána 10 , Saki Chan 9 , María Muñoz-Amatriaín 11 , Rachid Ounit 12 , Steve Wanamaker 11 , Thomas Schmutzer 1 , Lala Aliyeva-Schnorr 1 , Stefano Grasso 13 , Jaakko Tanskanen 14 , Dharanya Sampath 15 , Darren Heavens 15 , Sujie Cao 16 , Brett Chapman 3 , Fei Dai 17 , Yong Han 17 ,HuaLi 16 , Xuan Li 16 , Chongyun Lin 16 , John K. McCooke 3 , Cong Tan 3 , Songbo Wang 16 , Shuya Yin 17 , Gaofeng Zhou 2 , Jesse A. Poland 18 , Matthew I. Bellgard 3 , Andreas Houben 1 , Jaroslav Doležel 10 , Sarah Ayling 15 , Stefano Lonardi 12 , Peter Langridge 19 , Gary J. Muehlbauer 5,20 , Paul Kersey 7 , Matthew D. Clark 15,21 , Mario Caccamo 15,22 , Alan H. Schulman 14 , Matthias Platzer 8 , Timothy J. Close 11 , Mats Hansson 23 , Guoping Zhang 17 , Ilka Braumann 24 , Chengdao Li 2,25,26 , Robbie Waugh 6,27 , Uwe Scholz 1 , Nils Stein 1,28 & Martin Mascher 1,29 1 Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, 06466 Seeland, Germany. 2 School of Veterinary and Life Sciences, Murdoch University, Murdoch, Western Australia 6150, Australia. 3 Centre for Comparative Genomics, Murdoch University, Murdoch, Western Australia 6150, Australia. 4 Australian Export Grains Innovation Centre, South Perth, Western Australia 6151, Australia. 5 Department of Agronomy and Plant Genetics, University of Minnesota, St Paul, Minnesota 55108, USA. 6 The James Hutton Institute, Dundee DD2 5DA, UK. 7 European Molecular Biology Laboratory—The European Bioinformatics Institute, Hinxton CB10 1SD, UK. 8 Leibniz Institute on Aging—Fritz Lipmann Institute (FLI), 07745 Jena, Germany. 9 BioNano Genomics Inc., San Diego, California 92121, USA. 10 Institute of Experimental Botany, Centre of the Region Haná for Biotechnological and Agricultural Research, 78371 Olomouc, Czech Republic. 11 Department of Botany & Plant Sciences, University of California, Riverside, Riverside, California 92521, USA. 12 Department of Computer Science and Engineering, University of California, Riverside, Riverside, California 92521, USA. 13 Department of Agricultural and Environmental Sciences, University of Udine, 33100 Udine, Italy. 14 Green Technology, Natural Resources Institute (Luke), Viikki Plant Science Centre, and Institute of Biotechnology, University of Helsinki, 00014 Helsinki, Finland. 15 Earlham Institute, Norwich NR4 7UH, UK. 16 BGI-Shenzhen, Shenzhen 518083, China. 17 College of Agriculture and Biotechnology, Zhejiang University, Hangzhou 310058, China. 18 Kansas State University, Wheat Genetics Resource Center, Department of Plant Pathology and Department of Agronomy, Manhattan, Kansas 66506, USA. 19 School of Agriculture, University of Adelaide, Urrbrae, South Australia 5064, Australia. 20 Department of Plant and Microbial Biology, University of Minnesota, St Paul, Minnesota 55108, USA. 21 School of Environmental Sciences, University of East Anglia, Norwich NR4 7UH, UK. 22 National Institute of Agricultural Botany, Cambridge CB3 0LE, UK. 23 Department of Biology, Lund University, 22362 Lund, Sweden. 24 Carlsberg Research Laboratory, 1799 Copenhagen, Denmark. 25 Department of Agriculture and Food, Government of Western Australia, South Perth, Western Australia 6150, Australia. 26 Hubei Collaborative Innovation Centre for Grain Industry, Yangtze University, Jingzhou, Hubei 434025, China. 27 School of Life Sciences, University of Dundee, Dundee DD2 5DA, UK. 28 School of Plant Biology, University of Western Australia, Crawley 6009, Australia. 29 German Centre for Integrative Biodiversity Research (iDiv) Halle-Jena-Leipzig, 04103 Leipzig, Germany. *These authors contributed equally to this work. www.nature.com/sdata/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sdata.2017.44 24