Full text
ORIGINAL ARTICLE Metaproteogenomic insights beyond bacterial response to naphthalene exposure and bio-stimulation Marı ´a-Eugenia Guazzaroni 1,9,11 , Florian-Alexander Herbst 2,9 ,Iva ´nLores 3,9 , Javier Tamames 4,9 , Ana Isabel Pela ´ez 3 , Nieves Lo ´pez-Corte ´s 1 , Marı ´a Alcaide 1 , Mercedes V Del Pozo 1 , Jose ´Marı ´a Vieites 1 , Martin von Bergen 2,5 , Jose ´Luis R Gallego 6 , Rafael Bargiela 1 , Arantxa Lo ´pez-Lo ´pez 7 , Dietmar H Pieper 8 , Ramo ´n Rossello ´-Mo ´ra 7 , Jesu ´sSa ´nchez 3,10 , Jana Seifert 2,10 and Manuel Ferrer 1,10 1 Department of Biocatalysis, Institute of Catalysis, CSIC, Madrid, Spain; 2 Department of Proteomics, UFZ-Helmholtz-Zentrum fu¨r Umweltforschung GmbH, Leipzig, Germany; 3 A ´rea de Microbiologı ´a, IUBA, Facultad de Medicina, Universidad de Oviedo, Oviedo, Spain; 4 CSIC, Centro Nacional de Biotecnologı ´a, Madrid, Spain; 5 Department of Metabolomics, UFZ-Helmholtz-Zentrum fu¨r Umweltforschung GmbH, Leipzig, Germany; 6 IUBA, A ´rea de Prospeccio ´n e Investigacio ´n Minera, Universidad de Oviedo, Mieres, Spain; 7 IMEDEA (CSIC-UIB), Esporles, Spain and 8 Helmholtz Zentrum fu¨r Infektionsforschung–HZI, Braunschweig, Germany Microbial metabolism in aromatic-contaminated environments has important ecological implications, and obtaining a complete understanding of this process remains a relevant goal. To understand the roles of biodiversity and aromatic-mediated genetic and metabolic rearrangements, we conducted ‘OMIC’ investigations in an anthropogenically influenced and polyaromatic hydrocarbon (PAH)-contaminated soil with (Nbs) or without (N) bio-stimulation with calcium ammonia nitrate, NH 4 NO 3 and KH 2 PO 4 and the commercial surfactant Iveysol, plus two naphthalene-enriched communities derived from both soils (CN2 and CN1, respectively). Using a metagenomic approach, a total of 52, 53, 14 and 12 distinct species (according to operational phylogenetic units (OPU) in our work equivalent to taxonomic species) were identified in the N, Nbs, CN1 and CN2 communities, respectively. Approximately 10 out of 95 distinct species and 238 out of 3293 clusters of orthologous groups (COGs) protein families identified were clearly stimulated under the assayed conditions, whereas only two species and 1465 COGs conformed to the common set in all of the mesocosms. Results indicated distinct biodegradation capabilities for the utilisation of potential growth-supporting aromatics, which results in bio-stimulated communities being extremely fit to naphthalene utilisation and non-stimulated communities exhibiting a greater metabolic window than previously predicted. On the basis of comparing protein expression profiles and metagenome data sets, inter-alia interactions among members were hypothesised. The utilisation of curated databases is discussed and used for first time to reconstruct ‘presumptive’ degradation networks for complex microbial communities. The ISME Journal (2013) 7, 122–136; doi:10.1038/ismej.2012.82; published online 26 July 2012 Subject Category: integrated genomics and post–genomics approaches in microbial ecology Keywords: biodiversity; bio-stimulation; label-free protein quantification; metagenomics; metaproteomics; polyaromatic hydrocarbon Introduction Polyaromatic hydrocarbons (PAHs) are widely distributed in the environment owing to their abundance in crude oil and their widespread use in chemical manufacturing (Ka ¨stner, 2000). PAHs are pollutants of great concern owing to their toxicity, mutagenicity and carcinogenicity. A number of microorganisms, via the action of Rieske non-haem iron oxygenases, have the ability to grow on PAHs as a sole carbon and energy source (Beil et al., 1998; Roling et al., 2003; Witzig et al.,2006;Penget al., 2008). The majority of efforts aimed at understanding microbial responses to aromatic compounds have Correspondence: J Sa ´nchez, A ´rea de Microbiologı ´a, IUBA, Facultad de Medicina, Universidad de Oviedo, Oviedo, Spain. E-mail: [email protected] or J Seifert, Department of Proteomics, UFZ-Helmholtz-Zentrum fu ¨r Umweltforschung GmbH, Leipzig, Germany. E-mail: [email protected] or M Ferrer, Department of Biocatalysis, Institute of Catalysis, CSIC, Marie Curie 2, Madrid 28049, Spain. E-mail: [email protected]s 9 These authors contributed equally to this work. 10 These authors contributed equally to this work. 11 Current address: Laboratory of Molecular Ecology, Centro de Astrobiologı ´a (CSIC-INTA), Carretera de Ajalvir km 4, Torrejo ´nde Ardoz 28850, Madrid, Spain. Received 9 January 2012; revised 11 June 2012; accepted 11 June 2012; published online 26 July 2012 The ISME Journal (2013) 7, 122–136 & 2013 International Society for Microbial Ecology All rights reserved 1751-7362/13 www.nature.com/ismej
been focused on genomic (Jime ´nez et al., 2002; Nogales et al., 2008; Pucha"ka et al., 2008), transcriptomic (Domı ´nguez-Cuevas et al., 2006; Yuste et al., 2006) and proteomic (Santos et al., 2004; Segura et al., 2005; Kim et al., 2006; Kurbatov et al., 2006; Toma ´s-Gallardo et al., 2006) analyses conducted in pure cultures. Among PAH-(including naphthalene, fluorene, phenanthrene, pyrene and dibenzofuran, to cite some) mineralising strains, a number of bacteria have received special attention. Those include members of the common genera of Pseudomonas,Sphingomonas,Cycloclasticus, Burkholderia,Rhodococcus,Polaromonas, some novel genera of Neptunomonas and Janibacter,some thermophilic bacteria of Nocardia and Bacillus, some anoxic bacteria of Deltaproteobacteria and Alcaligenes, and high molecular weight PAHdegrading bacteria of Mycobacterium,Stenotrophomonas and Pasteurella; a list with known pure bacterial cultures capable of degrading PAHs can be seen in a recent review by Lu et al. (2011). However, a large number of culture-independent techniques have shown that pollutant-degrading organismsenrichedinthelaboratoryoftendonot have an important role in the in situ biodegradation of pollutants and that the diversity of pollutant-degrading organisms in polluted environmental samples is much greater than the diversity determined from cultivation (Abulencia et al., 2006; Liu et al., 2009; Boronin and Kosheleva, 2010; Yagi et al., 2010; Liang et al., 2011). Under this scenario, it is relevant to use molecular microbial tools to define key catabolic players at contaminated sites to predict pollutant degradation networks in the environment and to suggest methods for rational intervention associated with the implementation of bioremediation (VilchezVargas et al., 2010). However, the number of integrative ‘omic’ investigations that have been carried out in PAH-associated microbial communities is limited (Kweon et al., 2007; Powell et al., 2008; Selesi et al., 2010), because of the incomplete genomic information and curated databases available (Pe ´rez-Pantoja et al., 2009, 2012). Taking all of the above information into consideration, we performed a thorough and holistic (or ecosystems biology approach) phylogenetic, functional and proteomic analysis of the key players in a sample of strongly anthropogenically influenced, PAH-contaminated soil (N) and a naphthaleneenriched community derived from this soil (CN1). This study was carried out using also samples of the same soil bio-stimulated with calcium ammonia nitrate, NH 4 NO 3 and KH 2 PO 4 , and the commercial surfactant Iveysol (Ivey International Inc., Campbell River, BC, Canada) (herein referred as Nbs), and the naphthalene-enriched community derived from it (CN2). In all cases, the biodegradation networks of the respective whole communities were reconstructed. This study clarified the genomic and proteomic basis for the purpose of understanding microbial biodiversity, ecology and function in response to both PAHs (represented by naphthalene) and bio-stimulation. It should be noted that our metaproteomic investigations were restricted to naphthalene-enriched communities because of their lower complexity as compared with the soil samples as well as the larger assembled sequences obtained through direct pyro-sequencing for both samples. Materials and methods General methods and ‘OMIC’ data analysis and processing Full descriptions of the materials and methods used for soil characterisation; hydrocarbon analysis; DNA extraction; construction of 16 S RNA gene clone libraries, sequencing and phylogenetic analysis; denaturing gradient gel electrophoresis analysis; metagenomic setup and sequencing, assembly and gene prediction; metaproteomic setup and protein extraction, separation and identification and data processing are available in the Supplementary Materials and methods. Soil sample collection and preparation of biostimulated soil and enrichment cultures Soil samples (herein referred to as N) were obtained from a parcel on the northern Iberian Peninsula contiguous to a chemical plant (Lugones, Oviedo, Spain; 4011803300N, 313903100W, at an altitude of 300 m). The chemical plant (Supplementary Figure S1) was used for several decades for the production of naphthalene, phenols and other compounds from coal tar processing as well as for manufacturing resins. A considerable amount of other chemical products (pesticides, solvents, etc.) was stored, although probably not manufactured, in the plant. In 1989 the plant was closed and then used for years to store chemical waste, particularly polychlorinated biphenyls, coolants and other unspecified products. In 2006 most of the buildings were demolished and the characterisation of soil contamination started. Thus, a significant amount of the PAH detected were conceivably formed by the enriched bacterial populations present in the soil through a natural attenuation effect. The top soil was sampled at a depth of 0–30 cm on February 2007 (soil temperature 18 1C). Three replicates (1 kg each) were collected within a 1 m distance, and the samples were kept in open plastic bags in the dark at 4 1C. Vegetation and other non-soil materials, including cobbles, were removed prior to homogenising the samples. Immediately after acquiring the soil samples, they were sieved (2-mm mesh size), followed by mixing of representative subsamples of the triplicates samples, and 100 g aliquots of sample N were used for chemical analysis. In addition, 10 g aliquots were stored at 20 1Cin sterile flasks for DNAand proteome-based analyses. Bioremediation was performed over approximately 900 m 3 of the contaminated soil (N), where Microbial response to polyaromatics M-E Guazzaroni et al 123 The ISME Journal
9 tons of dehydrated calcium ammonia nitrate and 3 tons of a mixture of dehydrated NH 4 NO 3 and KH 2 PO 4 combined with 4500 l of the commercial surfactant Iveysol were applied to the homogenised soil. The C:N:P ratio that was applied was 100:10: 1, as recommended for bioremediation purposes (Gallego et al., 2007a). The treatment was performed over 231 days, and the biopile was routinely watered and tilled to maintain humidity (15–20%) and adequate oxygenation. Samples of the top bio-stimulated soil, herein named Nbs, were taken and processed using the same protocols as described above for the original soil. Enrichment cultures were performed in Bushnell Haas (Sigma Chemical Co., St Louis, MO, USA) mineral medium that contained naphthalene at a concentration of 0.1% (w/v) as the sole carbon source as described previously (Gallego et al., 2007b). The composition of the medium was as follows: MgSO 4 H 2 O (0.20 g l 1 ), CaCl 2 2H 2 O (0.02 g l 1 ), KH 2 PO 4 (1.0 g l 1 ), K 2 HPO 4 (1.0 g l 1 ), NH 4 NO 3 (1.0 g l 1 ) and FeCl 3 (0.05 g l 1 ). Two different inocula were used. The CN1 enrichment culture was obtained by inoculating 1 g of the polluted soil (N) into a flask containing 100 ml of the culture medium; the CN2 culture was inoculated with 1 g of the same soil that had been subjected to a massive bioremediation (bio-stimulation) process (Nbs). The enrichment cultures were incubated at 30 1C and 250 r.p.m., in which 0.1% (v/v) of the culture was transferred to fresh medium each week. Results and discussion Sample characteristics An agronomic analysis of the soil N revealed a loamy clay soil with a pH of 8.2 and a conductivity of 0.13 dS m 1 clearly containing low amounts of the typically predominant ions (calcium, magnesium, potassium and sodium); the detected natural organic matter, nitrogen and phosphorus levels in the soil (Supplementary Table S1) are characteristic of infertile soils. Gas chromatography–mass spectrometry (GC–MS) was used to quantify the levels of the 16 EPA (Environmental Protection Agency) priority PAHs present in the soil (Supplementary Figure S2). Totally, the soil exhibited an average concentration of 805 mg PAH per kg (Table 1), which is in the range or slightly lower to the level found in previously reported PAH contaminated soils (that is, 1667 mg kg 1 in Nı ´ Chadhain et al. (2006); 589 mg kg 1 in Richardson et al. (2012); 335–8645 mg kg 1 in Thavamani et al. (2012); 3000 mg kg 1 in Lors et al. (2012)). The biostimulated soil exhibited a total concentration of 221.6 mg kg 1 (average) of the 16 EPA priority PAHs and the quantified compounds are listed in Table 1. The degradation rate of naphthalene in N was too low and was difficult to be established owing to the long time transcurred in the presence of the contaminants; however, once submitted to biostimulation, we calculated the degradation rate of naphthalene in soil from time 0 to 231 days: 1.818 mg kg 1 soil per day. Denaturing gradient gel electrophoresis was used to survey the development of the microbial communities in the enrichment cultures and to deduce the time point at which a stable community was obtained. We observed that the denaturing gradient gel electrophoresis profiles for CN1 enrichment cultures changed between the 45 and 60 transfers (Supplementary Figure S3); changes in the profiles were also observed for CN2 after 35 and 60 transfers. These changes were subsequently maintained and samples subjected to at least 60 transfers were herein used for further investigations. The major bands were excised and sequenced, and their affiliation revealed the presence of microorganisms identified as members of the genera Achromobacter, Flavobacterium and Acidovorax in both enrichments, and Pseudomonas,Microbacterium,Lysobacter and to an endosymbiont of Acanthomaeba in CN2. Results indicated that consortia CN1 and CN2 were able to degrade 97% naphthalene at an initial concentration of 1000 p.p.m. within 72 and 80 h, respectively. Bacterial diversity and composition blueprints DNA isolated from each investigated microbial community was employed for a PCR-based 16S recombinant DNA (rDNA) gene diversity survey of the community structures in the samples. For this purpose, clone libraries were constructed as Table 1 Level of 16 EPA priority PAHs present in the original (N) and bio-stimulated (Nbs) soils PAH Concentration (mg kg 1 ) Soil N Soil Nbs Naphthalene 607 174 Acenaphthylene 1.60 0.40 Acenaphthene 16.60 4.40 Fluorene 20.80 5.96 Anthracene 21.20 4.99 Phenanthrene 46.40 24.23 Fluoranthene 30.40 4.73 Pyrene 19.40 4.02 Benzo(a)anthracene 12.00 2.83 Crysene 12.50 2.64 Benzo(b)fluoranthene 4.40 1.34 Benzo(k)fluoranthene 3.10 0.90 Benzo(a)pyrene 4.90 1.27 Indene(1,2,3-cd)pyrene 3.40 1.58 Benzo(g,h,i)perylene 1.50 0.82 Dibenzo(a,h)anthracene 0.79 0.18 Abbreviations: N, non-bio-stimulated polyaromatic hydrocarboncontaminated soil; Nbs, bio-stimulated polyaromatic hydrocarboncontaminated soil; PAH, polyaromatic hydrocarbon. Gas chromatography–mass spectrometry (GC–MS) was used for quantification. Microbial response to polyaromatics M-E Guazzaroni et al 124 The ISME Journal
described in the Supplementary Materials and methods, and the clones were fully sequenced to obtain the almost complete cloned 16S rDNA gene sequence. Additionally, and because of the low yields in the clone libraries of Nbs, we used the 16S rDNA gene partial sequences obtained in the metagenome survey. For this purpose, we used only those sequences that had a length 4600 nucleotides. All shorter sequences were discarded. A total of 670 sequences (that is, N: 212; Nbs: 261 (86 clones þ175 454 partial sequences); CN1: 90; CN2: 107) were obtained, and analysed. The overall phylogenetic composition in the libraries is shown in Figures 1–3, in which all operational phylogenetic units (OPUs) affiliated with twelve phyla of the domain Bacteria, namely, Proteobacteria followed to much lower extend by Bacteroidetes,Verrucomicrobia,Firmicutes,Actinobacteria,Chloroflexi,Cyanobacteria, Planctomycetes,Spirochaetes and the candidatus phyla OP8, TM7 and WCHB1 (Figure 4). The clones were classified into 124 bacterial operational phylogenic units (OPUs) (Lo ´pez-Lo ´pez et al., 2010) (Table 2 and Supplementary Figure S2). Higher values for Shannon–Wiener’s and Good’s coverage indexes indicated that communities N and Nbs were much more complex (51 OPUs and 53 OPUs were detected, respectively) than the CN1 and CN2 (13 OPUs and 12 OPUs, respectively) communities, which were dominated by few microbial species or strains and exhibited a rather simple structure (Figures 1–3). It should be recalled that each of the distinct OPUs detected was identified as a putative single species owing to the high sequence identity (Yarza et al., 2008). However, it should also be noted that the sequencing survey covered over 75% of the expected OPU diversity in all cases as it can be deducted from the high Good’s coverage in all samples (Table 2), and the rarefraction curves indicating closeness to saturation (Supplementary Figure S4). A large proportion of the obtained sequences affiliated with known families and clustered with branches represented by cultured microorganisms that have previously been found in contaminated environments (Figures 1–3 and Supplementary Figures S5 and S6), and thus, they do not appear to be specific to the soil and the enrichment samples that were investigated here. Whatever the case, despite soil characteristics and specific pollutants composition may differ, the biodiversity found in the original soils (N and Figure 1 A neighbour-joining tree of the proteobacterial 16S rDNA gene sequences from the gene sequences originated from N, Nbs, CN1 and CN2 communities. For the reconstruction, only the almost complete sequences of the clones from the four samples had been used. Besides, partial 454 sequences of 4600 nucleotides were inserted in the tree using the Parsimony tool implemented in ARB. The number of sequences comprised in each identity cluster (OPU) is specified for each sample. The colour code for the 16S rDNA sequences is as follows: N (green), Nbs (purple), CN1 (blue), CN2 (red). Whether an OPU contains sequences of more than one sample, the OPU denomination is written in black followed by the code colour and sequence number of each studied sample. Microbial response to polyaromatics M-E Guazzaroni et al 125 The ISME Journal
Nbs) (Shannon indexes of 2.81 and 3.32, respectively) are in the range of that found in other PAHcontaminated soils. For example, Shannon indexes of 1.74–2.78 (Vivas et al., 2008), 2.38–2.96 (Thavamani et al., 2012) and 4.87–5.01 (Martin et al., 2012) were observed in soils that had been historically contaminated with PAHs. The dominant group among the Gammaproteobacteria consisted of members of the genus Pseudomonas (Figures 2 and 4). The N sequences were distributed rather evenly within this genus, with approximately 40% of the total gammaproteobacterial sequences branching in the P. fluorescens, P. frederiksbergensis and P. amygdale lineages, followed by those affiliated with the Pseudoxanthomonas spp. (20%) and P. putida (14%) lineages (Supplementary Table S2). A similar even distribution was observed in the sample Nbs, being as well the members of Pseudomonadaceae the most abundant. However, in this case, the Pseudomonas species composition shifted towards the major representatives P. tuomourensis,P. pertucinogena, P. stutzeri genomovar 1, and yet undescribed Pseudomonas species. These results clearly contrast with the distribution of the CN2 sequences, where just representatives of P. stutzeri (46%) lineages (Supplementary Table S2) were observed in accordance with this being one of the major represented lineages in Nbs. Finally, in CN1, no representation of the Pseudomonas genus was found, but six sequences representing Pseudoxanthomonas were identified (6% of the total) (Figures 1 and 4 and Supplementary Table S2); this is in agreement with previous observations made in PAH microcosm experiments that suggest that, although, Pseudomonas spp. can be easily enriched and use a variety of organic substrates, they may be repressed by other Proteobacteria (Rhee et al., 2004). The alphaproteobacterial sequences detected in N were widely distributed among Rhizobium, Devosia,Novosphingobium,Sphingomonas and Brevundimonas, which were the most represented ;;)1VG1UPO (Pseudomonas stutzeri 1N 12Nsb 48CN2 sbN3)(2UPO Pseudomonas alcaligenes Pseudomonas otitidis, AY953147 ;)3VG(3UPO Pseudomonas stutzeri 6N 1Nsb ,).ps(4UPO Pseudomonas stutzeri 8Nbs 1CN2 ;)C41M(5UPO Pseudomonas stutzeri 5N 10Nbs Nbs132 (OPU 96) bsN12)(6UPO Pseudomonas tuomuerensis Pseudomonas mendocina, Z76674 N-179 (OPU 7) Pseudomonas argentinensis, AY691188 Pseudomonas flavescens, U01916 Nbs117 (OPU 97) sbN8)(8UPO Pseudomonas pertucinogena OPU 9 ( sp.) 20NbsPseudomonas )(01UPO Pseudomonas guineae 1Nbs4N, Pseudomonas anguilliseptica, X99540 Pseudomonas borbori, AM114527 N56)(11UPO Pseudomonas amygdali N-193 (OPU 12) Pseudomonas sp. VET-8, EU781734 Pseudomonas brassicacearum, AF100321 Pseudomonas fluorescens, AJ308308 Pseudomonas orientalis, AF064457 N22)(31UPO Pseudomonas putida Pseudomonas japonica, AB126621 OPU 14 sp.) 3Nbs(Pseudomonas Azotobacter vinelandii, AB175657 N-05 (OPU 15) Pseudomonas balearica, U26418 uncultured bacterium, GQ979955 OPU 16 ( sp.) 3NbsPseudomonas 10% N clones CN1(N) clones CN2 (Nbs) clones Nbs clones and 454 sequences Figure 2 A subset of the tree shown in Figure 1, where the genealogical composition of the members of the family Pseudomonadaceae is shown. Layout, reconstruction and colour codes used are the same as in Figure 1. Microbial response to polyaromatics M-E Guazzaroni et al 126 The ISME Journal
genera (Figure 1 and Supplementary Table S2). On the other hand, members of this lineage were barely represented in Nbs with just seven sequences, being Brevundimonas the most represented. None of these genera were detected in any of the enrichment cultures (Figure 1). In contrast, most of the alphaproteobacterial clones in CN1 (73%) were affiliated with Azospirillum species, that is, Azospirillum oryzae, whereas only three alphaproteobacterial sequences (out of a total of 108) were found in CN2, and only one was affiliated with Azospirillum (Figure 5 and Supplementary Table S2). Finally, the two N clones (out of 214) representing Betaproteobacteria affiliated with Acidovorax and Achromobacter spp. (Supplementary Table S2). On the other hand, Nbs sequences were importantly represented in this lineage with 27 sequences (10%), affiliated with Tetrathiobacter and Acidovorax.In this regard, the percentage of Betaproteobacteria found in the enrichment cultures was much higher (Figures 1 and 4); we observed an overrepresentation of sequences affiliated with the genera Comamonas and Achromobacter in CN1 and Achromobacter in CN2 (Figures 1 and 5). In addition to Delta-andEpsilon-proteobacteria, members of Tenericutes, Verrucomicrobia, Firmicutes and Cyanobacteria were only detected in samples N and Nbs (Figures 3 and 4). The above data demonstrated that the CN1 and CN2 communities displayed considerably different phylogenetic structures (Figures 4 and 5), sharing only two OPUs: Betaand Gamma-proteobacteria representatives of the denitrifying Achromobacter (that is, Achromobacter spanius) and Azospirillum (that is, Azospirillum doebereinerae) genera, respectively. Whereas the first one was most abundant in CN2 (29 vs 7 sequences), the second was particularly enriched in CN1 (35 vs 1 sequences) (Supplementary Table S2). In addition to supplying biosurfactants, the addition of different nitrate compounds may have stimulated the growth of denitrifying Achromobacter spp. and P. stutzeri (Song and Ward, 2003) in the bio-stimulated community, whereas depletion of nitrate may have favoured the presence of the nitrogen-fixing Azospirillum species in the non-stimulated community. In addition to denitrification, P. stutzeri is well known as naphthalene degrader, which has been isolated in such hydrocarbon enrichments from a wide range of environmental samples (Rossello ´- Mora et al., 1994), and was already an important component of the original Nbs soil. However, the possibility that the selective desorption and liberation of recalcitrant compounds from soils during the bio-stimulation process may have key functions in shaping the genomes and community structures of the associated microorganisms should not be ruled out. The findings of previous studies using BTEX biodegradation co-cultures (Shim et al., 2005) agree with this hypothesis. In any case, although bacteria related to Pseudomonas, Achromobacter Figure 3 A neighbour-joining tree of non-proteobacterial 16S rDNA gene sequences from the clone libraries established from N, Nbs, CN1 and CN2 community DNA. Layout, reconstruction and colour codes used are the same as in Figure 1. Microbial response to polyaromatics M-E Guazzaroni et al 127 The ISME Journal
and Comamonas spp. have been widely associated with PAH degradation in a number of contaminated ecosystems and microbial consortia (Goyal and Zylstra, 1997), the presence of the Azospirillum species in contaminated soils and PAH-degrading bacterial consortia has only previously been reported in a heavily creosote-contaminated soil (Vin ˜as et al., 2005). This is of special interest because it is known that the recently sequenced Azospirillum sp. B510 harbours gene sets for degradation pathways, though their functions have not yet been experimentally elucidated (Kaneko et al., 2010). Due to the clear shifts in community composition detected among the investigated conditions, it is important to establish the roles and degradation capabilities of the individual microbial members of the communities to understand the overall function of each community and its members. Genomic signatures associated with the aerobic degradation of aromatics DNA isolated from each microbial community was sequenced using a Roche GS FLX DNA sequencer, at Lifesequencing SL (Valencia, Spain) which produced 1 471 821 reads (N: 165 791; Nbs: 438 821; CN1: 448 713; CN2: 418 293) with an average length per read of 395 bp for CN1 and CN2, 336 bp for N and 531 for Nbs. Accordingly, a total of 55.7, 233.3, 177.2 and 165.2 Mbp of raw DNA sequences for N, Nbs, CN1 and CN2 were obtained, which were assembled into 2.6 Mbp (4335 contigs), 17.9 Mbp (16 032 contigs), 20.0 Mbp (20 809 contigs) and 13.0 Mbp (9915 contigs), respectively. CN1 (266) and CN2 (256) contained a higher number of contigs with lengths greater than 10 kbp compared with Nbs (68) and to major extent N (only 1); this may be due Figure 4 Comparison of bacterial phylotypes (cut-off of 497% sequence identity) based on the 16 S rDNA sequences extracted from N, Nbs, CN1 and CN2 community DNA. The percentages of bacterial phylogenetic lineages detected were based on OPUs. (a) Percentage of 16S rDNA sequences distributed by phyla. (b) Contribution of dominant groups among Gammaproteobacteria.(c) Relative distribution of major bacterial phylotypes (based on 16S rDNA gene sequences (OPUs)) for soil communities CN1 and CN2 at the genus level. Table 2 Statistical indexes for the four samples N Nbs CN1 CN2 Taxa (OPUs) 51 53 13 11 Individuals (sequences) 212 261 90 107 Dominance_D 0.13 0.06 0.20 0.30 Shannon_H 2.81 3.32 2.02 1.50 Good’s_G 0.76 0.80 0.86 0.90 Abbreviations: CN1, naphthalene-enriched community derived from N soil; CN2, naphthalene-enriched community derived from Nbs soil; OPU, operational phylogenetic unit; N, non-bio-stimulated polyaromatic hydrocarbon-contaminated soil; Nbs, bio-stimulated polyaromatic hydrocarbon-contaminated soil. PAST software v1.82b (Hammer et al., 2001) was used to calculate the statistical indices for the bacterial sequences. The following formulas were used. Shannon–Weiner index: H ¼S(ni/nt) Ln (ni/nt), where ni is the number of sequences of a particular OTU and nt is the total number of sequences. Dominance_D: D ¼S(ni/nt)2. Good’s coverage: C¼1(ni/nt), where ni is the number of OTUs observed exactly once and nt is the total number of sequences. Microbial response to polyaromatics M-E Guazzaroni et al 128 The ISME Journal
to the high diversity of the N sample, which resulted in a large pool of reads that could not be assembled (Supplementary Figure S7). Further information regarding the obtained metagenome characteristics is given in Supplementary Table S3. With a threshold of higher than 95% identity and an aligned length of more than 150 bp, 91.7% (for N), 95.2% (for Nbs), 92.3% (for CN1) and 89.9% (for CN2) of the predicted genes were assigned to particular genera, the analysis of which showed results comparable to those found in the 16S rDNA assignments (see Figure 5 for CN1 and CN2 comparison). Approximately 89% of the predicted genes (63 974 in total) were assigned to COG protein families, and 65% were assigned to KEGG pathways (Supplementary Table S3). Among the 1538 (for N), 2751 (for Nbs), 2775 (for CN1) and 2507 (for CN2) KEGG orthologs detected, 891 were shared between the two enrichment cultures, and 66% of these were putatively attributed to Achromobacter species. We first calculated the overrepresentation of functions between metagenomes; for that, we applied a z-test for independent proportions proposed by Li (2009) to statistically analyse the changes of the functional categories between samples. It should be noticed that the comparison was restricted to samples Nbs, CN1 and CN2 because the lower number of pyro sequences and COGs in sample N may have impact on the comparative result. We found a rather stable distribution of the functional categories composition between the samples with significant different contribution of only 238 out of 3293 COGs (Supplementary Table S4). Among them, 56/23 did show a significant increase in the bio-stimulated soil as compared with the CN1/CN2 enrichments, 58/26 increased in CN1 as compared with Nbs/CN2 and 41/55 were overrepresented in CN2 as compared with Nbs/CN1. Overall, we observed that both bio-stimulated communities were particularly enriched in COGs related to ‘Replication and repair’ and ‘Translation’ (29 vs 4 distinct COGs in total in CN1). Overrepresentation of such functions is typical for microbial communities developed under very dynamic environmental conditions, such as bio-stimulation; the addition of different nitrate compounds may have stimulated a competition between fast-growing organisms and organisms capable of metabolising poly-aromatics. By contrast, it is noteworthy that COGs related to ‘Transcription’ were most likely characteristics of enrichment cultures (13 and 7 COG enriched in CN1 and CN2, respectively), whereas only two (COG0789 and COG1974) were enriched in Nbs (Supplementary Table S4). This suggests a plausible scenario in which the stress caused by aromatic compounds during the naphthalene cultivation led to an enrichment of genomes containing transcriptional elements that could be required for stress endurance (Domı ´nguez-Cuevas et al., 2006) and/or activation of functions required for substrate uptake. Ten distinct COGs within the ‘Cell wall, membrane, envelop biogenesis’ category were found overrepresented in CN1, a number much higher that those found in Nbs (four COGs) and in CN2 (two COGs) (Supplementary Table S4). In addition, CN1 was also particularly enriched in ‘ABC Transport Systems’ by meaning of genes distributed in 11 distinct overrepresented COGs, for which only three and two were found in Nbs and CN2, respectively. COG0318 acetyl-CoA synthethase was also associated to CN1 community; enzymes of the COG0318 catalyse the formation of acetyl-CoA from acetate, suggesting that acetate may be a major end product used to produce acetyl-CoA feeding the Krebs pathway in members of CN1 (in agreement with proteomic data that will be discussed below). By contrast, CN2 was particularly enriched in genes coding proteins Figure 5 Relative abundances and distributions of gene coding sequences in the CN1 and CN2 communities based on taxonomic bins for bacterial 16S rDNA gene sequences (OPUs) (blue), taxonomic bins for metagenome-derived genes encoding proteins that could be assigned a taxonomic annotation (red) and taxonomic categories of proteins that were identified in the metaproteomes (green). As shown, the metagenome and the 16S rDNA analyses yielded congruent results. Further, only a minor number of bacterial members seemed to be metabolically active in both communities, by meaning of expressed proteins binned to them. Microbial response to polyaromatics M-E Guazzaroni et al 129 The ISME Journal
containing Fe–S clusters, namely COG3210 (large exo-proteins involved in haem utilisation or adhesion), COG1049 (aconitase B), COG0348 (polyferredoxin) as well as COG4773 (outer membrane receptor for ferric coprogen and ferric-rhodotorulic acid) (Supplementary Table S4), indicating iron being particularly relevant for CN2 community members. Taken together, the most notable observations drawn from all these data sets are that bio-stimulated communities appear to be extensive reservoirs of genes involved in replication, repair and translation and iron adhesion and utilisation whereas bacterial components of non-stimulated communities seems to be most active in cell wall, membrane and envelop biogenesis; in addition, enrichment most likely favoured transcriptional events. The obtained coverage was sufficient to produce a further partial theoretical metabolic reconstruction (including both exclusive and common capabilities) of the aerobic aromatic catabolic routes in the investigated communities; although these reconstructions were incomplete and likely represent composite cell networks, the information obtained may be sufficient to achieve a better understanding of how the four communities behave regarding their biodegradation capabilities on a genomic scale. To do that, a proper functional assignment of the predicted genes was performed using an in-house database containing protein sequences with biochemical functions shown to be involved in biodegradation (for details see Supplementary Materials and methods). According to this protocol, 428 (N: 38; Nbs: 115; CN1: 132; CN2: 143) open-reading-frame fragments were identified as showing close sequence similarity to genes that encode enzymes known to be involved in the aerobic metabolism of aromatics via diand trihydroxylated intermediates and, accordingly, presumptive functions were assigned. The overall gene features and presumptive catabolic routes most likely associated to them, are shown in Figure 6 Potential aerobic degradation networks of aromatics via diand tri-hydroxylated intermediates in the four investigated communities. The colour code used for the respective pathways is as follows: black, all samples; green, N and Nbs (specific for soil communities); blue, CN1 (N) (specific for non bio-stimulated communities]; red, CN2 (Nbs) (specific for bio-stimulated communities). Illustrations were created by ChemDraw graphic programme Chem Draw Ultra 8.0 (CambridgeSoft; http://www.cambridgesoft.com/) based on substrate specificity of enzymes listed in Supplementary Table S5, and the corresponding metabolic pathways were established based on bibliographic records. It should be noticed that the identification of a particular activity in CN1 and CN2 implies its presence in N and Nbs, respectively, although their signatures were not identified in the later most likely owing to their higher biodiversity and low coverage of the metagenome. Accordingly, those activities were indicated as CN1 (N) and CN2 (Nbs). Microbial response to polyaromatics M-E Guazzaroni et al 130 The ISME Journal