scieee AI-readable full text Open interactive document viewer

Mass Spectrometry-Based Proteomic Analysis of Selected Bacteria from the Human Gut Microbiome

Genth, Jerome

Abstract

Human gut bacteria live in a dynamic environment, constantly adapting their proteomes to changes like pH and nutrient availability. This study investigates how selected members of the human gut microbiome (HGM) respond to such conditions using mass spectrometry-based proteomics. Initially, the study evaluated two quantification methods—bottom-up label-free quantification (LFQ) and tandem mass tag (TMT)—in Bacteroides thetaiotaomicron. Both methods showed comparable results, indicating that this bacterium primarily alters the abundance of proteins involved in the machinery required to utilize the provided carbon sources The LFQ approach was then applied to examine proteomic changes in B. thetaiotaomicron, Blautia producta, and Bifidobacterium longum in response to different environmental pH levels. Distinct and concurrent alterations in pathway-related and stress-associated proteins were identified, including histidine biosynthesis in B. producta, nitrogen metabolism in B. thetaiotaomicron, and inositol carbohydrate metabolism in B. thetaiotaomicron and B. producta. The third project focused on identifying novel proteins, specifically short open reading frame-encoded peptides (SEP), in B. producta. The combined bottom-up and top-down proteomics analyses identified a total of 45 SEP, including previously reported SEP (BP1 to BP14). Their production varied based on environmental factors like media, pH, and supplements. The last project optimized a protocol for isolating extracellular vesicles (OMVs) from E. coli. A proteoform-directed top-down analysis, including a discovery-based open modification search, identified several potential post-translational modifications.

Full text

Mass Spectrometry-Based Proteomic Analysis of Selected Bacteria from the Human Gut Microbiome Dissertation zur Erlangung des Doktorgrades der Mathematisch-Naturwissenschaftlichen Fakultät der Christian-Albrechts-Universität zu Kiel vorgelegt von Jerome Genth Kiel, 2024 Erster Gutachter (First reviewer): Prof. Dr. Andreas Tholey Zweiter Gutachter (Second reviewer): Prof. Dr. Ruth Schmitz-Streit Tag der Disputation (Date of defense): 16.10.2024 Zum Druck genehmigt (Approved for publication): 16.10.2024 „Who you are is defined by what you're willing to struggle for.” - Mark Manson P A G E | i ABSTRACT Human gut bacteria live in a highly dynamic environment with frequent changes such as variations in pH or nutrient availability. To ensure their survival and functionality under these fluctuating conditions, they constantly adapt their proteomes by adjusting the abundance of essential proteins. However, the specific proteomic changes in individual human gut microbiome (HGM) members under different in vitro conditions remain largely unexplored. This study employs mass spectrometry-based proteomic analysis to address this gap on the selected members of the human gut microbiome. Initially, the reproducibility and accuracy of bottom-up label-free quantification (LFQ) and tandem mass tag (TMT)-based quantification were assessed in quantifying proteomic changes of Bacteroides thetaiotaomicron induced by sucrose and glucose. Both methods achieved comparable results, indicating that this bacterium primarily alters the abundance of proteins involved in the machinery required to utilize the provided carbon sources. The LFQ approach was then applied to examine proteomic changes in B. thetaiotaomicron, Blautia producta, and Bifidobacterium longum in response to different environmental pH levels. Distinct and concurrent alterations in pathway-related and stress-associated proteins were identified, including histidine biosynthesis in B. producta, nitrogen metabolism in B. thetaiotaomicron, and inositol carbohydrate metabolism in B. thetaiotaomicron and B. producta. The comparative analysis of bottom-up and top-down proteomics in B. producta demonstrated their ability to quantify the same proteins, whether differentially abundant or not. The third project focused on identifying novel proteins, specifically short open reading frameencoded peptides (SEP), in B. producta. A proteogenomic approach was used to analyze their presence under various cultivation conditions, including different media (BHI and YCFA), pH levels, and supplemented factors (yeast extract, SCFAs, and LPS). Stringent validation criteria ensured accurate identification of SEP, and biochemical predictions explored their potential functional roles. The combined bottom-up and top-down proteomics analyses identified a total of 45 SEP, including previously reported SEP (BP1 to BP14). This study demonstrated that the production of certain SEP in B. producta is influenced by specific environmental factors rather than solely by interspecies interactions, as previously suggested. The last project established and optimized an isolation protocol for the LC-MS-based proteomic analysis of extracellular vesicles. Bottom-up proteomic analysis of outer membrane vesicles (OMVs) isolated from E. coli under different culture conditions and growth phases quantified OMV marker proteins and classified the subcellular topological distribution of the P A G E | ii E. coli OMV proteomes. Nanoparticle tracking analysis validated OMV nanoparticles, and an in vitro wound healing assay with human colonic Caco-2 cells suggested a dose-dependent inhibitory effect of OMVs on wound closure. Various sample preparation protocols were evaluated, and the integration of non-ionic detergents for OMV lysis and proteolytic digestion improved protein and peptide identifications, making the improved in-solution digestion protocol the preferred method for future analyses. In most projects, a proteoform-directed top-down analysis, including a discovery-based open modification search, was applied. This analysis identified several potential post-translational modifications (e.g., β-methylthio-aspartic acid on B. thetaiotaomicron ribosomal protein S12 and seryl-phosphorylations on B. producta HPr proteins) and neo-termini (potential artificial cleavage events, initiator methionine excisions, and alternative initiation sites). P A G E | ix TABLE OF CONTENTS Abstract ..................................................................................................................................... i Zusammenfassung .................................................................................................................. iii Danksagung ............................................................................................................................. v List of Publications and Conference Contributions ................................................................. vii Table of Contents .................................................................................................................... ix I GENERAL INTRODUCTION .................................................................................................. 1 II GENERAL METHODS ....................................................................................................... 17 III PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON ........................................................... 37 IV INFLUENCE OF PH ON BACTERIAL PROTEOMES ............................................................... 63 V PROTEOGENOMIC ANALYSIS OF B. PRODUCTA .............................................................. 103 VI OUTER MEMBRANE VESICLES ANALYSIS ....................................................................... 121 Bibliography ............................................................................................................................... I List of Abbreviations ........................................................................................................... XXIII List of Figures ..................................................................................................................... XXV List of Tables ................................................................................................................... XXVIII Appendix ............................................................................................................................ XXIX P A G E | 1 I GENERAL INTRODUCTION 1 The Dynamic Nature of the Proteome .............................................................................. 2 1.1 Proteome Variation .................................................................................................... 3 2 Principles of Mass Spectrometry ..................................................................................... 5 2.1 Online LC-MS ............................................................................................................ 5 2.2 Mass-to-charge Analysis ........................................................................................... 5 2.3 Tandem Mass Spectrometry and Fragmentation Analysis ........................................ 7 3 Analysis of Mass Spectrometry Data ............................................................................... 8 3.1 Data Acquisition Techniques ..................................................................................... 9 3.2 Peptidoform and Protein Identification .................................................................... 10 3.3 Quantitative Techniques in Proteomics ................................................................... 11 4 Inflammatory Bowel Disease .......................................................................................... 12 4.1 The Human Gut Microbiome as a Therapeutic Target ............................................ 12 4.2 Proteomics in IBD Research ................................................................................... 14 5 Objective of the Thesis ................................................................................................... 15 I | GENERAL INTRODUCTION P A G E | 2 1 The Dynamic Nature of the Proteome The genome contains all genetic information and typically remains largely stable within a specific cell. In contrast, the proteome, representing the complete set of proteins (Wasinger et al., 1995), is constantly changing. This constant alteration of the proteomic landscape is characterized by continuous protein synthesis, degradation, and modification. The ability of cells to adjust their protein composition continuously in response to changing needs and influences, both internal and external, is a key aspect of cellular flexibility and crucial for a cell to effectively carry out its functions. The fundamental process of transcribing DNA (deoxyribonucleic acid) into mRNA (messenger ribonucleic acid) molecules and then translating them into amino acid sequences results in a wide variety of proteins. Each protein is assembled from the repertoire of 22 proteinogenic amino acids, with each amino acid imparting unique physico-chemical characteristics to the resulting protein. The term 'protein' is essentially a generic term referring to a canonical amino acid sequence (Cassidy et al., 2023). With an increased understanding of the proteome, it becomes evident that the traditional notion of 'one-gene, one-protein, one-function' is inadequate to fully comprehend the complexity of the proteome (Carbonara et al., 2021; Cassidy et al., 2023). Recognizing this complexity, terms such as 'protein species' (Schlüter et al., 2009) and 'proteoforms' (Smith and Kelleher, 2013) have emerged, with 'proteoforms' gaining widespread acceptance (Carbonara et al., 2021; Marx, 2024). The proteoform concept includes a broad spectrum of molecular variations resulting from coand post-translational modifications (PTMs) and sequence variants, which go beyond transcription and translation errors (FIGURE I-1). Understanding the function of diverse proteoforms is essential for comprehending cellular dynamics and processes (Marx, 2024). FIGURE I-1 | Sources of Proteome Complexity. Isoform variation of the same gene combined with sitespecific changes generate a variety of proteoforms. Abbreviation: SAV (single amino acid variation). Adapted with permission from (Aebersold et al., 2018). DNA Isoforms Site-specific features Proteoforms SAV Glycosylation Phosphorylation GENERAL INTRODUCTION | I P A G E | 3 1.1 Proteome Variation Organisms have developed many effective strategies to diversify the proteome without increasing the size of the genome, using mechanisms that generate multiple proteins from a single gene. Geneand Transcript-level Variation – Transcriptional read-through, a fundamental mechanism, allows RNA polymerases to bypass termination signals and transcribe multiple operons within a single mRNA molecule (Wade and Grainger, 2014). In response to environmental cues, bacteria dynamically exhibit transcriptional read-through to adapt to changing conditions (Junier and Rivoire, 2016). In eukaryotes, transcriptional complexity arises not only from transcript elongation but also from mRNA splicing, during which internal exons are removed. Variations in transcripts, such as ribosomal frameshifting, can lead to translation initiation and termination at different sites within a single mRNA molecule (Atkins et al., 2016), resulting in polypeptides with sequences differing from the main open reading frame (ORF) (Korniy et al., 2019). In particular, alternative translation initiation and synthesis of N-terminally truncated polypeptides are associated with the development of certain diseases (Bogaert et al., 2020). Additionally, RNA editing, involving processes such as nucleotide deamination of adenosine into inosine or nucleotide insertions and deletions, can alter mRNA sequences and lead to the formation of distinct mRNA isoforms (Knoop, 2011). Translation-level Variation – Coand post-translational modifications can rapidly modify protein properties and functions. Currently, over 200 types of PTMs or biological and chemical modifications, primarily targeting specific amino acid residues (FIGURE I-2), have been documented (Creasy and Cottrell, 2004; Montecchi-Palazzi et al., 2008). For example, phosphorylation and lipidation, can redirect proteins to specific cellular locations or modulate their interactions with other molecules, thereby significantly influencing various physiological processes such as transcription, translation, and metabolic functions (Jiang et al., 2018; Macek et al., 2019). Additionally, other PTMs such as glycosylation, can influence protein folding and stability (Jayaprakash and Surolia, 2017). Controlled proteolysis, by endopeptidases, can precisely cleave proteins, converting inactive zymogens into their biologically active forms (Neurath and Walsh, 1976). Zymogens typically contain inhibitory propeptides at their Ntermini, ensuring both the protection of protein function and their directed transport to specific cellular compartments. Upon reaching their destination, the propeptide is removed, activating the catalytic activities of the protein. For example, cathepsins, a family of lysosomal proteases, achieve their optimal functionality within the highly acidic environment of lysosomes upon activation (Jordans et al., 2009). This selective, compartment-specific activation isolates potentially detrimental reactions, thereby safeguarding the cell. In contrast, exopeptidases that I | GENERAL INTRODUCTION P A G E | 4 primarily remove terminal amino acids play a significant role in protein degradation processes. Both proteolytic processes possess the potential to generate truncated proteoform variants, capable of modulating protein activities and potentially influencing the development of specific diseases (Bogaert et al., 2020). The analysis of PTMs requires precise mass spectrometry (MS) analysis because isobaric PTMs share similar molecular masses and, consequently, comparable mass-to-charge ratios. This similarity increases the risk of misidentifications due to measuring errors (Kim et al., 2016). For example, trimethylation (C3H6, 42.047 Da) and acetylation (C2H2O, 42.011 Da) on lysine residues, commonly found on histone proteins, have remarkably similar masses, differing by only 0.036 Da. In the case of a typical 1 kDa tryptic peptide (Fricker, 2015), this corresponds to a difference of 36 ppm (parts per million). While modern high-resolution MS mass analyzers have mitigated this problem for peptides, challenges remain for intact proteins. For a 15 kDa core histone, this mass difference is 2.4 ppm, highlighting the critical role of accurate mass measurements in determining the correct proteoform. FIGURE I-2 | Protein Modifications in Bacteria. Overview of the most commonly occurring protein posttranslational modifications, corresponding amino acid residue, and corresponding modifying enzymes. Reactive groups on amino acid side chains are highlighted. PTS, phosphotransferase system; TCS, two-component system. Adapted with permission from (Macek et al., 2019). GENERAL INTRODUCTION | I P A G E | 5 2 Principles of Mass Spectrometry Mass spectrometry is a powerful analytical technique used for identifying molecules based on their mass-to-charge ratio (m/z). 2.1 Online LC-MS The preferred method in proteomics involves the online coupling of a separation technique, such as liquid chromatography (LC), with MS. Peptide and protein mixtures are typically separated based on their hydrophobic interactions with a non-polar stationary phase, employing reversed-phase high-performance liquid chromatography (RP HPLC). Commonly used microparticulate columns consist of silica beads with covalently bonded hydrophobic chains, such as octadecyl alkane chains (C18) for peptides and butyl alkane chains (C4) for protein mixtures. This results in species with differing hydrophobicity eluting at different times due to their affinities for the stationary phase. In contrast, monolithic columns feature a continuous porous structure for the separation of peptides or proteins. Retention, a metric that quantifies how long an analyte remains on the column, and resolution, the ability to distinguish adjacent peaks, are determined by the balance between the analyte's hydrophobic interactions with the stationary phase and its solubility in the mobile phase. The gradual increase in the concentration of organic solvent, typically acetonitrile, in the mobile phase initiates elution. The utilization of an acidic mobile phase, often containing an ion-pairing modifier such as formic acid or TFA, can enhance chromatographic separation, aid in controlling retention times, improve peak shapes, and improve the overall detection efficiency and performance of the LCMS system (García, 2005; Lenčo et al., 2022). Electrospray ionization (ESI) is a widely used technique for transferring analytes from the liquid phase to the gas phase in online LC-MS setups. It enables the gentle vaporization of molecules without causing significant fragmentation (Kebarle and Verkerk, 2009). During the ESI process, the analyte solution is directed through a conductive capillary (referred to as the emitter), to which a voltage is applied. In positive ion mode, the preferred polarity for the analysis of proteins and peptides, ions are converted into protonated molecular ions [M+nH]n+ with various charge and protonation states. Protonation primarily occurs at the free amino terminus and the basic side-chain functionalities of arginine, lysine, and histidine residues. 2.2 Mass-to-charge Analysis Various mass analyzers and detectors, each with unique characteristics such as mass accuracy, speed, sensitivity, and resolution (the ability to distinguish closely spaced m/z ratios), have been developed for measuring ion m/z ratios. High-resolution and accurate mass analyzers allow for the identification of characteristic isotope patterns, particularly in peptides I | GENERAL INTRODUCTION P A G E | 6 containing naturally occurring stable, heavy isotopes such as 13C and 15N. Mass analyzers operate on diverse principles: some (e.g., linear or quadrupole ion traps) utilize different electric currents to manipulate ions for selective isolation, trapping, and fragmentation, while others (e.g., time-of-flight or Orbitrap analyzers) accelerate ions towards a detector in an electric field to measure their time-of-flight or radial oscillation frequencies to determine m/z (Savaryn et al., 2016). Today, hybrid instruments combine different analyzers allowing flexibility in experimental design. A notable example of such hybridization is the Fusion Lumos Tribrid mass spectrometer, integrating quadrupole, Orbitrap, and dual-pressure linear ion trap analyzers (FIGURE I-3). Before entering the mass spectrometer, ions can be introduced into a high-field asymmetric waveform ion mobility spectrometry (FAIMS) module, to separate them based on ion mobility variations in the presence of high and low electric fields (Guevremont, 2004). Acting as a mass filter, FAIMS can eliminate singly charged contaminant ions and enhance the depth of proteome coverage by applying multiple compensation voltages (CVs) (Swearingen and Moritz, 2012; Kaulich et al., 2022a). After entering the Fusion Lumos Tribrid mass spectrometer, ions pass through ion optics, which focus and direct them toward the quadrupole. The quadrupole then acts as a mass filter, allowing to select precursor ions of interest based on their m/z values (Savaryn et al., 2016). Ions within a specific isolation window are collected, stored in the ion-routing multipole, and then transferred through the C-trap into the Orbitrap. Fragment spectra can be acquired in either the Orbitrap or the ion trap mass analyzer. FIGURE I-3 | The Orbitrap Fusion Lumos Tribrid Mass Spectrometer. After ionization, ion optics focus and direct gas phase analyte ions toward the quadrupole. In the quadrupole, ions are selectively filtered based on their mass-to-charge ratio (m/z) within a specific isolation window. The selected precursor ions are then directed to the ion-routing multipole or the linear ion trap mass analyzer for fragmentation. Fragment ion spectra are then acquired in either the Orbitrap or the linear ion trap. Adapted from http://planetorbitrap.com. Ion Source Transfer Tube Ion Optics Quadrupole Orbitrap C-Trap Ion Routing Multipole Performs HCD Dual-Pressure Linear Ion Trap Performs CID and ETD GENERAL INTRODUCTION | I P A G E | 7 2.3 Tandem Mass Spectrometry and Fragmentation Analysis Analyzing samples via tandem mass spectrometry (MS/MS) with data-dependent acquisition (DDA) involves obtaining precursor MS1 spectra, fragmenting isolated ions, and determining the m/z values of resulting fragment ions in the MS2 spectra. This process provides essential insights into the amino acid sequence of the peptide (Biemann, 1992), facilitating the differentiation of isobaric peptides with distinct amino acid sequences or PTMs (Kim et al., 2016). Tribrid instruments offer various ion activation methods, with collision-induced dissociation (CID) (Hunt et al., 1986) or higher-energy collisional dissociation (HCD) (Olsen et al., 2007) being the most common. CID and HCD involve collisions with inert gases (e.g., nitrogen, helium, or argon), resulting in charge-directed fragmentation, driven by the kinetic energy of ionizing protons, which mostly results in cleavage of the peptide backbone. In the Fusion Lumos, CID is performed in the high-pressure collision cell of the linear ion trap, while HCD is conducted separately in the ion routing multipole. This separation allows for enhanced resolution and improved mass accuracy measurements of fragment ions by enabling higher kinetic energies and shorter impact times during fragmentation (Michalski et al., 2012). Another important method is the EThcD approach, combining electron transfer dissociation (ETD) (Syka et al., 2004) in the linear ion trap with subsequent transfer of precursors and product ions to the collision cell for HCD fragmentation. Fragmentation of precursor ions by ETD results in radical-driven fragmentation of peptide bonds while limiting the neutral loss of labile groups and preserving PTMs (Syka et al., 2004). Fragment ions resulting from peptide dissociation are influenced by several factors, including the mass and charge state of the precursor ion, amino acid composition, and adjacent residues at the backbone cleavage site (Reid et al., 2001; Tabb et al., 2003; Haverland et al., 2017). Typically, the charge is distributed between both fragments, with fragment ions retaining their charge either at the C-terminus (x-, y-, and z-ions) or the N-terminus (a-, b-, and c-ions) according to the Roepstorff-Fohlmann-Biemann nomenclature (Roepstorff and Fohlman, 1984; Biemann, 1992) (FIGURE I-4). FIGURE I-4 | Fragment Ion Nomenclature. (A) Roepstorff–Fohlmann–Biemann nomenclature for fragment ions (Roepstorff and Fohlman, 1984; Biemann, 1992). (B) Determination of peptide sequence using y-ion series. y11 z11 x11 b2c2 a2 y2z2 x2 b11 c11 a11 a2-ion b2-ion c2-ion x2-ion z2-ion y2-ion A B y1 y2 y3 y10 y9 y8 y7 y6 y5 y4 y12 y11 R A V E A D H P T D A N 200 1,2001,0008006004000 m/z Relative intensity (%) 100 50 0 TNADTPHDAEVAR I | GENERAL INTRODUCTION P A G E | 8 3 Analysis of Mass Spectrometry Data Unlike genomics and transcriptomics, which can amplify DNA or RNA using techniques like polymerase chain reaction, respectively, proteomics lacks a comparable method for amplifying proteins. This limitation necessitates careful attention to the proteomics workflow to maximize information from the potentially limited starting material. In general, a proteomics workflow can be divided into three major steps (FIGURE I-5). The initial step (i) involves sample preparation, which includes isolating the protein mixture from the studied biological sample. The complexity of samples can be reduced and proteome coverage enhanced by employing protein or peptide pre-fractionation techniques (Zhang et al., 2010). Following sample cleanup, the second step (ii) involves online separation and MS measurement of analytes, typically using reversed-phase LC-MS/MS. In principle, peptide or protein masses can be determined from MS1 spectra by utilizing the m/z value and spacing between isotope peaks. The corresponding amino acid sequences are then determined based on the MS2 mass difference between the fragments. With modern mass spectrometers generating thousands of spectra in a short time, manual annotation becomes impractical, necessitating an automated approach for peptide and protein identification. Thus, the final step (iii) involves employing various software applications and computational tools to identify peptides and proteins, along with performing post-processing tasks like quantification, statistical testing, and metabolic mapping to evaluate significant differences between samples. In this study, Proteome Discoverer served as the primary software tool for proteomic data analysis. Notably, Proteome Discoverer demonstrated improved performance compared to the widely used software MaxQuant, particularly in terms of quantification yield, dynamic range, and reproducibility for label-free quantification (Palomba et al., 2021). FIGURE I-5 | Generalized Bottom-up Proteomics Workflow. Proteins are extracted from cells treated with various stimuli. Sample complexity can be reduced through protein prefractionation, peptide fractionation, or isobaric labeling after enzymatic digestion. Following sample clean-up using solidphase extraction, samples are separated by reversed-phase liquid chromatography (LC) and analyzed by mass spectrometry (MS). Analytes are identified from MS/MS spectra using Proteome Discoverer (PD) database matching. Post-processing includes quantification, statistical testing, and metabolic mapping to assess significant differences between samples. (Adapted with permission from (Mergner and Kuster, 2022)). m / z m / z m / z PD SAMPLE PREPARATION Cell culture Cell Lysis LC-MS/MS DATA ANALYSIS Protein digestion Sample cleanup Identification Database search Postprocessing Optional Protein fractionation Peptide fractionation Isobaric labeling MS/MS MS Theoretical spectrum Acquired spectrum Quantification Difference Significance GENERAL INTRODUCTION | I P A G E | 15 Although studies on single bacterial isolates in response to diverse in vitro conditions are limited, they have yielded diverse proteomic findings. For instance, proteomic analyses were utilized to validate B. thetaiotaomicron-specific protein requirements under various growth conditions (Liu et al., 2021a), observe the selective packing of acidic glycosidases and proteases into Bacteroides outer membrane vesicles (Elhenawy et al., 2014), and characterize proteins required for mucin degradation by human gut bacteria (Crouch et al., 2020). However, how proteomic profiles change in single members of the microbiota at IBD initiation, progression, and during therapeutic interventions remains largely unknown. To improve functional knowledge of the gut microbiome in the context of IBD, it is essential to generate detailed proteomic profiles. 5 Objective of the Thesis This study aims to analyze the proteomic adaptations of selected bacteria of the human gut microbiome in response to varying environmental conditions, focusing on the specific objectives: Evaluation of Quantitative Proteomic Approaches (Chapter III) – Assess the reproducibility and accuracy of bottom-up label-free quantification (LFQ) and tandem mass tag (TMT)-based quantification methods for detecting proteomic changes induced by sucrose and glucose in B. thetaiotaomicron. Exploration of pH-Dependent Proteomic Alterations (Chapter IV) – Analyze proteomic changes in B. thetaiotaomicron, B. producta, and B. longum in response to different environmental pH levels. Identification of Short Open Reading Frame-Encoded Peptides (Chapter V) – Employ a proteogenomic approach to identify and analyze the translation of previously undiscovered SEP in B. producta under various cultivation conditions. Optimization of OMV Isolation Protocol (Chapter VI) – Establish and optimize a protocol for the LC-MS-based proteomic analysis of extracellular vesicles, specifically outer membrane vesicles (OMVs) isolated from E. coli. Integration of Proteoform-Directed Analysis (Chapters III to IV) – Apply a proteoformdirected top-down analysis, including a discovery-based open modification search, to identify potential post-translational modifications and neo-termini in HGM members. P A G E | 16 P A G E | 17 II GENERAL METHODS 1 General Materials ............................................................................................................. 18 1.1 Chemicals and Reagents ........................................................................................ 18 1.2 LC-column Packing ................................................................................................. 18 2 Protein and Peptide Sources .......................................................................................... 19 2.1 Cultivation of Human Gut Bacteria .......................................................................... 19 2.2 Cultivation of E. coli and OMV Isolation .................................................................. 20 2.3 Caco-2 Wound Healing Assay ................................................................................ 22 3 Sample Preparation for LC-MS Analysis ....................................................................... 23 3.1 Proteome Clean-up and Digestion .......................................................................... 23 3.2 Molecular Weight-based Prefractionation ............................................................... 25 3.3 SDC Removal by Phase-transfer ............................................................................ 26 3.4 Solid-phase Extraction ............................................................................................ 27 3.5 Tandem Mass Tag Labelling ................................................................................... 27 3.6 Protein Gel and Staining ......................................................................................... 29 4 Mass Spectrometry Data Acquisition ............................................................................ 30 4.1 Bottom-up LC-MS Measurements ........................................................................... 30 4.2 Top-down LC-MS Measurements ........................................................................... 31 4.3 MALDI-TOF Measurements .................................................................................... 32 5 Data Processing and Analysis ....................................................................................... 32 5.1 Database Search ..................................................................................................... 32 5.2 Genome and Proteome Sequence-based Predictions ............................................ 34 5.3 Functional in silico Analysis ..................................................................................... 34 5.4 Functional and Statistical Data Analyses ................................................................ 35 II | GENERAL METHODS P A G E | 18 1 General Materials 1.1 Chemicals and Reagents Deionized water (18.2 MΩ/cm) was obtained using an Arium611 VF system (Sartorius). Complete protease inhibitor cocktail was purchased from Roche Diagnostics. Pierce Coomassie and BCA protein assay kits (both Thermo) were used for protein quantification. Single-pot, solid-phase-enhanced (SP3) bead-based purification was performed using SeraMag SpeedBead carboxylate-modified magnetic particles (GE Life Sciences). Sequencing grade modified trypsin was purchased from Promega. Lyophilized protein standards of cytochrome C (Equus caballus), myoglobin (Equus caballus), beta-casein (Bos taurus), carbonic anhydrase (Bos taurus), bovine serum albumin (Bos taurus), and alcohol dehydrogenase (Saccharomyces cerevisiae) were purchased from Sigma-Aldrich. HeLa and cytochrome C digest were purchased from Thermo Fisher. Synthetic peptides were purchased from JPT Peptide Technologies GmbH. Additional chemicals required for various stages of media preparation, sample handling, and LC-MS/MS analysis were purchased from various suppliers, including Sigma-Aldrich, Merck, and Serva. 1.2 LC-column Packing Frits for column packing were prepared by cross-linking Kasil (potassium silicate solution) with formamide (Cortes et al., 1987). Kasil 1, a 29.1% (w/w) potassium silicate solution in water, was solubilized at 80°C and 2000 rpm for 1 hour. Equal parts of Kasil 1 and a 25% w/v formamide solution in water were combined to create the final frit solution. After mixing and centrifuging at 21,000 g for 5 minutes at 20°C, tub fused silica capillaries (360 µm OD x 75 or 150 µm ID) were gently pressed onto a glass microfiber filter (GF/C, Whatman), which had been previously soaked with 2 µL of the frit solution (Maiolica et al., 2005). Capillaries were incubated at 85°C for 20 hours for polymerization. Using a high-pressure bomb loader, capillaries were packed with PLRP-S beads (5 µm, 1000 Å), which were collected from a PLRP-S column (4.6 x 50 mm, Agilent). The collected beads were washed with 50% and then 100% acetonitrile (ACN), dried at 70°C and suspended in methanol (60 mg/ml). After allowing the material to settle by gravity for 20-30 minutes and mixing for 1 minute and sonication, columns were packed under low-speed stirring (400-500 rpm) with continuous mechanical tapping (Kovalchuk et al., 2019). The bomb containing the capillary was pressurized to 100 bar with nitrogen immediately after mounting the capillary to prevent passive filling of the capillary with the solvent. After the packing process was completed, the pressure was slowly released for 10 minutes to prevent bubble formation within the column. The freshly packed columns were connected to an Ultimate 3000 system and flushed with 95% ACN at a flow rate of 600 GENERAL METHODS | II P A G E | 19 nl/min for 30 minutes to compress the sorbent bed. The columns were then cut to the desired length, and connections were made using ZIRCOFIT UHPLC fittings with a 1/16" 13-mm bore (MS Wil). This resulted in two different types of columns: pre-columns (150 µm x 4 cm) and analytical columns (75 µm x 17 cm). 2 Protein and Peptide Sources Details on the number of bacterial cell culture replicate and specific culture conditions for particular experiments are described in the experimental procedure sections of the corresponding chapters. 2.1 Cultivation of Human Gut Bacteria The cultivation of Bacteroides thetaiotaomicron VPI-5482, Blautia producta ATCC 27340, and Bifidobacterium longum NCC 2705 was performed by Kathrin Schäfer (Department of Infectious Diseases and Microbiology, University of Lübeck, UKSH Lübeck; chair: Prof. Dr. Jan Rupp). Yeast extract, casein, and fatty acid (YCFA) medium (Duncan et al., 2009) (TABLE II-1) or brain-heart infusion (BHI) medium adjusted to different pH values (pH 6.0, pH 7.0, and pH 8.0) and supplemented with different carbon sources (27.8 mM glucose or sucrose) or various growth supplements (lipopolysaccharide (LPS), SCFA or yeast extract) were utilized for bacterial cultivation. Bacterial cells were cultured at 37°C under strict anaerobic conditions in an anoxic chamber (H35, Don Whitley Scientific Limited) containing 85% (v/v) N2, 10% (v/v) CO2 and 5% (v/v) H2. A single colony was transferred to 5 ml of YCFA or BHI medium, incubated overnight and 0.1% (v/v) was used for inoculum. TABLE II-1 | Modified YCFA Medium COMPONENTS [g/l] COMPONENTS [mg/l] COMPONENTS [g/l] Casitone 10 Biotin 0,02 Acetic acid 1900 Yeast extract 2,5 Folic acid 0,10 Propionic acid 700 Carbon Source 5 Pyridoxine hydrochloride 0,05 iso-Butyric acid 90 MgSO4 x 7 H2O 0,45 Thiamine-HCl x 2 H2O 0,05 n-Valeric acid 100 CaCl2 x 2 H2O 0,90 Riboflavin 0,05 iso-Valeric acid 100 K2HPO4 0,45 Nicotinic acid 0,05 KH2PO4 0,45 D-Calcium pantothenate 0,001 NaCl 0,90 Vitamin B12 0,05 Resazurin 0,01 p-Aminobenzoic acid 0,05 Distilled water 4 Lipoic acid 10 NaHCO3 1 L-Cysteine HCl 0,1 Hemin 0,02 II | GENERAL METHODS P A G E | 20 2.2 Cultivation of E. coli and OMV Isolation Growth Media – Lysogeny broth (LB) medium (1% tryptone, 0.5% yeast extract, 1% NaCl, pH 7.0, Sigma-Aldrich) was prepared according to the manufacturer's instructions. Agar plates were prepared by adding 1% agar and autoclaving at 121°C for 15 minutes. M9 culture media were prepared as indicated in TABLE II-2 and TABLE II-3 and sterilized by autoclaving at 121°C for 12 minutes. Glucose and acetate solutions were autoclaved separately, and heat-sensitive components such as biotin and thiamine were filter-sterilized (0.2 μm filter) and added to the media afterward. TABLE II-2 | 100X Trace Elements Solution COMPONENT [g/l] CONCENTRATION EDTA 5 g/l 13.4 mM FeCl3-6H20 0.83 g/l 3.1 mM ZnCl2 84 mg/l 0.62 mM CuCl2-2H20 13 mg/l 76 µM CoCl2-2H20 10 mg/l 42 µM H3BO3 10 mg/l 162 µM MnCl2-4H20 1.6 mg/l 8.1 µM TABLE II-3 | M9 Mineral Medium VOLUME COMPONENT COMPONENT CONCENTRATION 100 ml M9 salt solution (10X) Na2HPO4-2H20 KH2PO4 NaCl NH4Cl 33.7 mM 22.0 mM 8.55 mM 9.35 mM 20 ml Carbon source Glucose or acetate 15 mM or 45 mM 1 ml MgS04 (1M) MgS04 1 mM 0.3 ml CaCl2 (1M) CaCl2 0.3 mM 1 ml* Biotin (1 mg/ml) Biotin 1 µg 1 ml* Thiamin (1 mg/ml) Thiamin 1 µg 10 ml* Trace elements (100X) Trace elements 1X * 0.22-µm filter sterilization Cultivation – A single colony of Escherichia coli K-12 strain MG1655 was cultured in 100 ml of LB medium for 18 h at 37°C and 150 rpm in a non-baffled 250 ml shake flask. Afterward, the culture medium was removed by centrifugation at 7,000 g for 3 min at 4°C, and cells were washed twice with filter-sterilized M9 minimal medium. Then, cells were inoculated with 1/100 dilutions of 18 h pre-culture at an initial OD600 of 0.1 in M9 media containing either 15 mM glucose or 45 mM acetate. The cultures were incubated aerobically at 37°C, 150 rpm in baffled 1 l shake flasks with 300 ml of media. During the cultivation culture samples were taken to track cell growth via OD600 measurements. GENERAL METHODS | II P A G E | 21 Isolation of Outer Membrane Vesicles (OMV) – OMV isolation from E. coli supernatant was performed using the ExoBacteria OMV Isolation Kit (System Biosciences). Cultivations were transferred into sterile 50 ml centrifuge tubes and centrifuged at 8,000 g for 20 min at 4°C. The supernatant was transferred to a new sterile 50 ml centrifuge tube and spun again at 8,000 g for 20 min at 4°C. The supernatant was filter-sterilized (0.2 μm filter) and the resulting cell-free supernatant was used to isolate OMVs. The binding column was prepared by adding 1 ml of OMV binding resin and equilibrating it with 10 ml of OMV binding buffer. After equilibration, the binding buffer was allowed to completely flow through the column. Subsequently, the bottom of the column was sealed, and 20 ml of supernatant was added. The top of the column was sealed, and the unit was placed on a rotating rack for 30 min at 4°C to allow for mixing and binding of the OMVs to the resin. After 30 min, the top and bottom of the column were opened to allow the supernatant to flow through the resin. Depending on the experiment, loading of supernatant was repeated up to 2 additional times so that a total of 40 ml or 60 ml of culture supernatant, was incubated with the OMV binding resin. After OMV binding, the supernatant was allowed to flow through the column, and the resin was washed with 15 ml OMV binding buffer per 20 ml of loaded supernatant. Afterward, the bottom of the column was sealed, and 1.5 ml OMV elution buffer was added. Columns were allowed to incubate at 20°C for 2 min with gentle agitation every 30 s, after which the bottom of the column was unsealed and the OMV isolate was collected. Samples were frozen at -80°C until further analysis. Nanoparticle Tracking Analysis – For nanoparticle tracking analysis (NTA) a NanoSight NS300 system (NanoSight Ltd), equipped with a 488 nm laser and a high sensitivity digital camera system sCMOS (scientific complementary metal oxide semiconductor camera) was employed. Videos were acquired and analyzed using the NTA software (version 3.3), with minimum track length and blur setting, all set to automatic. The camera shutter was set manually in dependency of the particle intensity and camera gain was set to 366. Camera levels were set to 14 to 15 and the detection threshold was set to 5, to reveal small particles. The ambient temperature was maintained at 25°C. Samples were administered and recorded under controlled flow, using the NanoSight syringe pump and script control system. For each sample, six videos of 60 seconds duration were recorded, with a 10-second delay between recordings, generating six replicate histograms that were averaged. Lastly, the hydrodynamic diameters and particle size distributions were analyzed by the software using the StokesEinstein equation. A summary of the complete list of parameters used in the nanoparticle tracking analysis is provided in TABLE II-4. The instrument was calibrated prior to each experimental run using standardized nanoparticle dilutions purchased from the manufacturer. II | GENERAL METHODS P A G E | 22 TABLE II-4 | Parameters for Nanoparticle Tracking Analysis PARAMETER SETTING Instrument NanoSight NS300 NTA Version NTA 3.3 Dev Build 3.3.301 Diluent Water (1:1) Camera Type sCMOS Camera Level 14-15 Laser Type Blue 488 Slider Shutter 1200 - 1260 Slider Gain 366 Isolated E. coli OMVs cultivated until midlogarithmic, prestationary, stationary or death phase were subjected to LC-MS/MS analysis to analyse the repertoire of OMV proteins which had been obtained from different growth phases, mid-logarithmic, prestationary and stationary for OMVs isolated from glucose and midlogarithmic, stationary and death phase for OMVs isolated from acetate. Isolated E. coli OMVs cultivated until midlogarithmic, prestationary, stationary or death phase were subjected to LC-MS/MS analysis to analyse the repertoire of OMV proteins which had been obtained from different growth phases, mid-logarithmic, prestationary and stationary for OMVs isolated from glucose and midlogarithmic, stationary and death phase for OMVs isolated from acetate. Isolated E. coli OMVs cultivated until midlogarithmic, prestationary, stationary or death phase were subjected to LC-MS/MS analysis to analyse the repertoire of OMV proteins which had been obtained from different growth phases, mid-logarithmic, prestationary and stationary for OMVs isolated from glucose and midlogarithmic, stationary and death phase for OMVs isolated from acetate. Isolated E. coli OMVs cultivated until midlogarithmic, prestationary, stationary or death phase were subjected to LC-MS/MS analysis to analyse the repertoire of OMV proteins which had been obtained from different growth phases, mid-logarithmic, prestationary and stationary for Frames/Sec 25 Frames 1498 Temperature 25.0 ºC Viscosity Water 0.889 cP Syringe Pump Speed 40 Detect Threshold 5 Blur Size Auto Max Jump Distance Auto 2.3 Caco-2 Wound Healing Assay Human Caco-2 cells were cultured in collaboration with Britta Steer (Systematic Proteome Research & Bioanalytics, Institute for Experimental Medicine, University of Kiel; chair: Prof. Dr. Andreas Tholey). Cells were cultured in Roswell Park Memorial Institute (RPMI) medium (Gibco-Invitrogen) supplemented with 10% fetal calf serum (FCS; Gibco-Invitrogen) and 1% penicillin/streptomycin in a 5% CO2, 95% humidity environment at 37°C. Cells were passaged weekly upon reaching 80% confluence. Wound healing assays were conducted using µ-dishes with inserts (Ibidi GmbM) at a density of 4×105 cells/cm2. After reaching a confluence of 7080%, cells were serum-starved overnight using serum-deprived medium (0.1% FCS), the insert was removed, and cell layers were washed twice with phosphate-buffered saline (PBS). Cells were then incubated with 2 ml of 0.1% FCS serum-deprived medium containing different concentrations of E. coli OMVs (10, 50, and 100 µg/ml), OMV elution buffer, LPS E. coli O55:B5 (1 µg/mL, Sigma-Aldrich), transforming growth factor β (TGFβ) (5 ng/mL, SigmaAldrich) and medium only. The migration process into the cell-free gap of approximately 500 µm was measured by taking microscopic photographs at 0, 12, and 30 hours using a digital camera on an inverted microscope at 10x magnification. The wound area was quantified using a wound healing plugin for ImageJ (Suarez-Arnedo et al., 2020). The percentage of wound closure was calculated relative to the medium control (arbitrarily assigned as 100%) based on the area measured immediately after insert removal (At=0h), as well as 12 and 30 hours after incubation (At=∆h) (eq. 1). Wound'closure% = /!!"#$"!!"∆$ !!"#$ 0×100% (eq. 1) GENERAL METHODS | II P A G E | 23 3 Sample Preparation for LC-MS Analysis Details regarding the quantities of protein and peptide used, the number of technical replicas, and different MS parameters for specific experiments are described in the experimental procedure sections of the corresponding chapters. Protein concentrations were determined in at least triplicate using the Pierce Coomassie or BCA Protein Assay Kit according to the manufacturer's instructions. 3.1 Proteome Clean-up and Digestion Human Gut Bacteria Samples – Cell lysis of human gut bacteria was performed by Kathrin Schäfer (Department of Infectious Diseases and Microbiology, University of Lübeck, UKSH Lübeck; Chair: Prof. Dr. Jan Rupp). Culture media was removed by centrifugation at 2,100 g for 10 min at 4°C, washed twice with water, and centrifuged again. Bacterial cells were suspended in lysis buffer (6 M GndHCl (guanidine hydrochloride), 100 mM HEPES (4-2hydroxyethyl-1-piperazineethanesulfonic acid), 20 mM NaCl and 1x cOmplete protease inhibitor, pH 7.5). Cells were lysed using ten cycles of freeze-thawing (30 s, -80°C in an ethanol bath followed by thawing in a sonication bath for 30 s). After centrifugation at 21,000 g for 20 min at 4°C, supernatants were collected and cell debris was washed twice with lysis buffer, centrifuged and supernatants were pooled. Disulfide bridges were reduced and alkylated using 12 mM tris(2-carboxyethyl)phosphine (TCEP) and 40 mM 2-chloroacetamide (CAA) for 1 hour at 25°C and 800 rpm. The samples were then precipitated using 9x volume ethanol at -20°C. After 16 hours of incubation at -20°C, the precipitates were centrifuged at 21,000 g for 10 min at 4°C and washed twice with cold ethanol. Residual ethanol was evaporated in a fume hood. Precipitates were suspended in 0.5 M GndHCl, 12.5 mM HEPES, or 100 mM triethylamoniumbicarbonat (TEAB) (all pH 8.5), digested by adding trypsin at a 1:40 enzyme to substrate ratio, and incubated overnight at 37°C on a shaker at 800 rpm. E. coli samples – 1 mg of E. coli cells were processed using the sample preparation by simple extraction and digestion (SPEED) method (Doellinger et al., 2020), with TFA added at a 1:4 (v/v) ratio of sample to TFA. Samples were incubated for 5 minutes at 20°C and neutralized with 2 M Tris base using 8-fold the volume of TFA used for lysis. Aliquots of 50 μg protein were reduced and alkylated by incubation in 10 mM TCEP and 40 mM CAA at 95°C for 5 minutes. Samples were diluted 1:5 with water and proteins were digested with trypsin at a 1:50 enzymeto-substrate ratio for 20 hours at 37°C. II | GENERAL METHODS P A G E | 24 OMV samples – 40 µg of OMV samples were processed using different sample preparation protocols, including in-solution, on-bead SP3 (Hughes et al., 2019), and a modified version of the on-membrane filter-assisted sample preparation (FASP) method (Manza et al., 2005; Wiśniewski et al., 2009). In-solution Digestion – OMV samples were lysed by ten cycles of freeze-thaw (30 s, -80°C in an ethanol bath followed by thawing in a sonication bath for 30 s) followed by adjustment to 100 mM TEAB using 1 M TEAB. Alternatively, 40 µg OMV samples were lyophilized before the addition of 50 µl of lysis buffer (2% (w/v) sodium deoxycholate (SDC), 0.1 M TEAB, pH 8.5). Samples were then incubated at 95°C for 5 minutes to inactivate proteases, reduced with 10 mM dithiothreitol (DTT) at 56°C for 1 hour, and alkylated with 50 mM iodoacetamide (IAA) in the dark at 20°C for 30 minutes. For samples containing SDC, the concentration was adjusted to 0.5% SDC with 0.1 M TEAB (pH 8.5) and trypsin was added at a 1:40 (w/w) enzyme-to-substrate ratio. The digestion reaction was performed overnight at 37°C on an orbital shaker at 800 rpm. On-bead Digestion – For OMV lysis, sodium dodecyl sulfate (SDS) or SDC was added to a final concentration of 1% or 2% (w/v), respectively. Subsequently, samples were incubated at 95°C for 5 minutes before being reduced using 10 mM DTT for 1 hour at 56°C and alkylated with 50 mM IAA in the dark for 30 minutes at 20°C. SP3 beads were prepared by mixing equal amounts of hydrophilic and hydrophobic beads, were washed twice with water, and resuspended in water to a final concentration of 20 μg/µl. For each sample, 20 µl of beads were added, and protein binding was induced by adding ethanol to a final concentration of 50% (v/v) and incubating for 5 minutes at 25°C at 800 rpm. The beads were immobilized using a magnetic rack, the supernatant was removed, and the beads were washed twice with 200 µl of 80% ethanol. Beads were resuspended in digestion buffer (100 mM TEAB, pH 8.5) containing trypsin at a 1:40 (w/w) enzyme-to-substrate ratio and incubated overnight at 37°C on a shaker at 800 rpm. Enhanced digestion was performed by adding 0.5% (w/v) SDC or 0.001% (w/v) dodecyl-β-D-maltosid (DDM) to the digestion buffer. The digests were acidified to pH 2 to 3 by adding trifluoroacetic acid (TFA), and the beads were pelleted by centrifugation. The supernatants were stored on ice until sample cleanup. For SDC digests, samples were additionally processed using a modified phase transfer protocol. On-membrane Digestion – For OMV lysis, a final concentration of 1% or 2% (w/v) of SDS or SDC, respectively, was added. The samples were incubated at 95°C for 5 minutes, reduced with 10 mM DTT for 1 hour at 56°C, and loaded onto Amicon centrifugal filter units (30K molecular weight cutoff, Millipore). After centrifugation at 12,000 g for 15 min at 20°C (all buffer exchanges were performed by centrifugation under identical conditions), SDS-lysed samples GENERAL METHODS | II P A G E | 31 adjusted according to gradient length (20 to 40 s), with ions of unassigned, +1, and >+8 charge states excluded, and lock mass (445.12003 m/z) enabled. 4.2 Top-down LC-MS Measurements Nanoflow LC-ESI-MS measurements were conducted using a Dionex Ultimate 3000 HPLC system online coupled to a Fusion Lumos Tribrid (Lumos) mass spectrometer (both Thermo Fisher Scientific). Depending on the applied method, the mass spectrometer was equipped with the field asymmetric ion mobility spectrometry (FAIMS) pro interface. Liquid Chromatography – Generally, 1 to 1.5 µg protein preparations were concentrated and washed onto a trap column (5 mM x 0.33 mm, 5 μm C4 resin, 300 Å; PepMap300, Thermo Fisher Scientific) for 5 min with 2% ACN and 0.05% aqueous TFA at a flow rate of 30 µl/min. Subsequently, proteins were separated on an analytical column (75 μm x 50 cm, 2.6 μm C4 resin, 150 Å; Accucore, Thermo Fisher Scientific) at 300 nl/min and separated within a 60-, 90or 140-min linear gradient of increasing LC solvent B (80% ACN, 0.1% FA) in LC solvent A (0.1% aqueous FA). Additionally, GELFrEE samples were separated on a self-packed PLRPS pre-column (150 µm x 4 cm, 5 µm PLRP-S resin, 1000 Å), before being separated on an analytical column (75 µm x 17 cm, 5 µm PLRP-S resin, 1000 Å). LMWP samples, GELFrEE fractions, and SCX fractions of B. thetaiotaomicron were separated using a 90-minute gradient from 15% to 55% B. In contrast, the LMWP samples of B. producta were separated using a 60-minute gradient from 15% to 60% B. The linear gradient was followed by a sharp increase to 98% solvent B for 2 min, an isocratic wash step for 13 min, and finally a column equilibration with 5% B for 15 min. Mass Spectrometry – A 15 Volt source-induced dissociation was applied to favor protein ion desolvation and the RF Lens was set at 30%. The acquisition was performed in “peptide mode”, according to the recommendations of vendors for the mass range below ca. 20−30 kDa. Dual CV FAIMS Method – Acquisition of MS spectra were obtained using two methods including two different FAIMS compensation voltages (CVs). The first method involved combining -60 V and -50 V, while the second method used -40 V and -30 V. Different microscan settings were used for each method, with only 2 microscans being utilized for -60 V, and 4 microscans being used for all other CVs. MS1 spectra were recorded in the Orbitrap from 500 to 1800 m/z, at a resolution of 120K, using an AGC target value of 200% and a maxIT of 246 ms. Both methods obtained MS2 spectra with a 60K resolution (3 s cycle time) after collision-induced dissociation (CID) with an NCE of 25%. MS2 spectra were acquired within 4 II | GENERAL METHODS P A G E | 32 microscans in a scan range of 500 m/z to 2000 m/z, using a normalized AGC target of 400% and a maxIT of 250 ms. Only precursors with a charge state between 4 and 50 or undetermined charge states were selected with enabled dynamic exclusion (exclude after 2 times, 60 s duration) plus and minus 2.5 m/z. Mutli-CV FAIMS Method – Four different CVs (-60, -50, -40, -20 V) with different MS1 and MS2 settings were applied, based on (Kaulich et al., 2022a). Within a cycle time of 3 s MS2 spectra were acquired with an isolation window of 5 m/z, 50K resolution, 400% AGC target, 250 ms injection time, and for fragmentation CID with an NCE of 30% was utilized. Only precursors with a charge state between 4 and 50 or undetermined charge states were selected with enabled dynamic exclusion (n= 2, 60 s). Settings for CVs -60 and -50 V: resolution 60K/50K (MS1/MS2), maximum injection time 118/125 ms, microscans 2/2, AGC target 200%/400%. Settings for CVs -40, -20 V: resolution 120k/60k, maximum injection time: 246/250 ms, 4/4 microscans. 4.3 MALDI-TOF Measurements Matrix-assisted laser desorption/ionization time of flight (MALDI-TOF) mass spectrometry was performed using a TOF/TOF 5800 mass spectrometer equipped with the 4000 series explorer software (both AB Sciex). The mass spectrometer accumulated 400 laser pulses at an intensity of 65%. Spectra were acquired within a mass range of 100 m/z to 1200 m/z. Prior to analysis, the instrument was calibrated using a six-peptide solution to calibrate the target m/z range. Samples were prepared by mixing with varying volumes (v/v) of a CHCA matrix solution (3 mg/ml α-cyano-4-hydroxycinnamic acid in 70% acetonitrile, 0.1% TFA) and spotted as 1 µL drops onto a 384-well Opti-TOF MALDI Insert (AB Sciex). 5 Data Processing and Analysis Details on the specific search parameters for each experiment and the version of Proteome Discoverer (PD) used are available in the experimental procedures section of each chapter. 5.1 Database Search Depending on the sample type, tandem mass spectra were searched and compared to the UniProt reference proteome of Bacteroides thetaiotaomicron vpi-5482, Blautia producta ATCC 27340, Bifidobacterium longum NCC 2705 or Escherichia coli K12 (UP000001414 / 4,782 proteins; UP000515789 / 5,372 proteins; UP000000439 / 1,725 proteins; UP000000625 / 4,450 proteins, accessed between September 2020 and August 2022) or the integrated GENERAL METHODS | II P A G E | 33 proteogenomics database (iPtgxDB) for Blautia producta ATCC 27340 (accessed November 2021) (Omasits et al., 2017). In addition, common repositories of adventitious proteins were included in the searches. Bottom-Up Searches – Peptide identification and quantification experiments were performed using PD (v.2.2, v.2.5, or v.3.0; Thermo Fisher Scientific) utilizing its built-in search engine SequestHT, INFERYS rescoring, or CHIMERYS algorithm. These included trypsin as the proteolytic enzyme with up to two or four missed cleavage sites allowed, carbamidomethylation of cysteine (57.021 Da) as fixed modification, oxidation of methionine (15.995 Da) as variable modifications, a precursor tolerance of 10 ppm and a fragment ion tolerance of 0.02 Da. For TMT experiments, TMT6 was employed as the fixed modification on lysine and peptide Ntermini. To assess under/over labeling, variable TMT6 modifications were added as search parameters for lysine and peptide N-termini, as well as fixed modifications for histidine, serine, threonine, and tyrosine. Additional searches, including phosphorylation as variable modifications on serine, histidine, threonine, tyrosine, or arginine (79.966 Da), and semi-tryptic searches, were customized based on the sample type. Retention-time alignment was applied for samples necessitating MS1 intensity-based quantification. All results were adjusted to 1% PSM, peptide, and protein FDR, employing a target-decoy approach using reversed protein sequences and posterior error calculation by the percolator algorithm (Käll et al., 2008). Top-Down Searches – Proteoform identification and quantification were performed using PD (v.2.5.0.400; Thermo Fisher Scientific) utilizing its built-in ProSightPD 4.1 or 4.2 nodes (Proteinaceous Inc.). These included the high/high cRAWler node in combination with Xtract for deconvolution. The annotated proteoform search, with a maximum of three proteoform spectrum matches (PrSMs) per precursor, a minimum of three matched fragments, and no delta M mode, was utilized to identify full-length proteoforms. Truncated proteoforms were detected using the subsequence search node, with a maximum of one PrSM per precursor and a requirement of at least six matched fragments. Searches, utilized a precursor and fragment mass tolerance of 10 ppm, with variable modifications acetylation (42.011 Da) and formylation (27.995 Da) at the proteoform N-terminus. Additional searches, specifically designed to search for fixed dehydro-modifications on cysteine (-1.008 Da) to identify potential disulfide bridges and open-modification searches with a 500 Da precursor window, were customized based on the sample type. For label-free quantification, raw data from multi-CV measurements were filtered using Freestyle v.1.6 based on FAIMS CVs, and the resulting filtered data was saved as distinct .raw files, effectively dividing them into four fractions. Quantification was carried out employing the High-Resolution Feature Detector node with the sliding window deconvolution algorithm, utilizing an average retention time width of 0.33 min II | GENERAL METHODS P A G E | 34 with mass features needed to be present in only 1% of the total data files. All results were subjected to an FDR correction for both PrSMs and proteoforms, with a threshold of 1%. 5.2 Genome and Proteome Sequence - based Predictions Protein sequences were retrieved from UniProt and genomic sequences were obtained from the European Nucleotide Archive (Leinonen et al., 2011). Unless otherwise noted, default parameters were used for all prediction tools. Protein sequences were annotated using WebMGA (Web Services for metagenomic analysis) (Wu et al., 2011) that used the COG (Clusters of Orthologous Genes) database for functional annotation (Galperin et al., 2021). For that, RPSBLAST was run on the prokaryotic NCBI COG database with an E-value cut-off of 0.001. Kyoto Encyclopedia of Genes and Genomes (KEGG) orthology and the prediction of pathway-specific metabolic functions, as well as the reconstruction of KEGG pathways, were facilitated through the utilization of BlastKOALA (KEGG Orthology And Links Annotation) (Kanehisa et al., 2016). This analysis was performed at the genus level using BLASTp to search and compare the data with a non-redundant dataset of pangenome sequences. Genomic sequences were subjected to subsystem annotation using the Rapid Annotation using Subsystem Technology (RAST) server and analyzed using the SEED viewer (Overbeek et al., 2014). A classic RAST (v. 2.0) annotation search along with FIGfam (release 70) was used for curation of the genomic data. As the automatic annotation process may run into problems, such as overlapping gene pairs or overlapping RNAs these errors and frameshifts were fixed automatically (even if that requires deleting some gene candidates). Debug statements were turned on and gaps were backfilled, allowing RAST to blast large gaps for missing genes. RAST facilitated the classification of genes into predefined subsystems, which represent groups of genes involved in specific biological processes or pathways. Potential protein-coding genes were cross-referenced to coding sequences (CDS) from the NCBI Prokaryotic Genome Annotation Pipeline (PGAP) (Tatusova et al., 2016) and linked back to UniProt accession. The integration of PATRIC (Pathosystems Resource Integration Center) was used to display and analyze RAST-annotated genomes to further investigate genome properties. 5.3 Functional in silico Analysis Physicochemical properties such as the isoelectric point (pI) and grand average of hydropathy (GRAVY) score were calculated using ProtParam with default settings for pK values (Gasteiger et al., 2005). Phobius was used for the prediction of protein localization using a posterior probability ≥ 0.5 (Käll et al., 2004). Potential antimicrobial peptide (AMP) activity and their GENERAL METHODS | II P A G E | 35 functional targets were assessed using AMPfun (a probability score of >0.5 indicates potential AMP activity, while a score <0.5 indicates non-AMP activity) (Chung et al., 2020). Disulfide bridges were predicted using SCRATCH (Cheng et al., 2005) and functional domains and motifs were predicted using NCBI´s Conserved Domains search (v. 3.20) (Wang et al., 2023a). Polysaccharide utilization loci (PULs) were assigned using the Polysaccharide Utilization Loci DataBase (PULDB) (Terrapon et al., 2015). The direction of enzymatic reactions was checked using the ExplorEnz enzyme database (https://www.enzyme-database.org/) (McDonald et al., 2007) and the MACiE database (Mechanism, Annotation, and Classification in Enzymes) (Holliday et al., 2005). 5.4 Functional and Statistical Data Analyses LFQ Data Normalization – Total protein and peptide concentrations were determined and normalized by BCA assay prior to LC-MS/MS. The data normalization pipeline consisted of the following steps: (I) Data cleanup by removing proteins from potential contaminants and those with low or medium confidence levels. (II) Total intensity normalization was performed by median intensity normalization. (III) The raw and normalized intensity data were tested for normal distribution and Pearson correlation using Matplotlib in Python. (IV) Removal of proteins identified in only one out of three or two out of five biological replicates. (V) Calculate the median of all biological replicates. TMT Data Normalization – Reporter ion intensities were corrected for signal interference by subtracting the percentage of interference from the measured reporter ion intensities (Savitski et al., 2013). For each PSM of the same TMT channel reporter ion intensities were normalized to one and the normalized median intensities were subtracted by the corresponding isolation interference. Occurring negative values due to recalculation were replaced by the minimum positive value in each channel. Only spectra with >50% isolation interference were used for relative quantification and subjected to normalization as described for LFQ. Differential Analyses and Statistical Rationale – The Perseus software suite (v.1.6.14.0) was utilized to perform functional 1D enrichment analyses, Fisher’s exact test, and two tails Welch's or Student's t-tests (Tyanova et al., 2016). Statistical tests were corrected for multiple testing applying a permutation-based or Benjamini-Hochberg FDR calculation based on the pvalue distribution at 1 or 5% (Benjamini and Hochberg, 1995). Differentially abundant proteins were identified with Log2 fold change ± 0.485. Significant differences were assessed using ANOVA with Dunnett's multiple comparison test to identify specific groups or conditions that were significantly different from a control group. Statistically significant differences were considered when * (p < 0.05); ** (p < 0.01); **** (p < 0.0001). Direction pathway analysis (DPA) II | GENERAL METHODS P A G E | 36 was performed using the directPA package in R (v.4.2.2) to perform test statistics in twodimensional space with a modified Pearson correlation test, to identify concordantly higher, lower, and discordantly abundant proteins (Yang et al., 2014). By specifying eight different directions, DPA was used to perform COG and KEGG pathway analysis on the selected directions, with a p-value ≤ 0.05 considered significant. Cleavage site specificity analysis was performed using iceLogo motifs via the standalone version of iceLogo (v.1.2) (Colaert et al., 2009). Prior to analysis, full tryptic peptides, peptides in which initiator N-terminal methionine excision, shared peptides, identical peptide sequences with multiple modifications, and peptides with canonical C-terminus were removed to eliminate false-positive cleavage sites. The calculated cleavage specificities were corrected for the natural abundance of the corresponding amino acids in the organism's proteome. Venn diagrams were created using Venny (Oliveros, 2007). The sequence coverage of top-down data was calculated using the protti package (Quast et al., 2022) in R (v.4.2.2). UpSet plots were generated using the UpSetplot and Matplotlib package (Lex et al., 2014) in Python (v. 3.11.1). Annotation of shotgun proteomics mass spectrometry data was performed using the Interactive Peptide Spectral Annotator (Brademan et al., 2019). Protein crystal structure predictions were generated using AlphaFold Colab v2.3.0 (Jumper et al., 2021) and visualized using PyMOL (Schrödinger, 2002). Experimental design workflows were created using BioRender.com. iBAQ Calculation – iBAQ values were obtained by dividing a protein’s total non-normalized intensity by the number of theoretically observable tryptic peptides between 7 and 30 amino acids with up to 2 missed cleavages (eq. 4). iBAQ =∑'intensity/#theoretical'peptides' (eq. 4) To obtain the relative iBAQ value (riBAQ) for each protein, the iBAQ value was normalized by dividing by the sum intensity of all iBAQ values (eq. 5). riBAQ =iBAQ/∑iBAQ (eq. 5) To estimate the relative abundance of histidine-containing proteins, the riBAQ values were scaled to 100% and multiplied by the absolute number of histidine residues in each protein (eq. 6). riBAQ(*+,)=riBAQ×#Histidine'residue' (eq. 6) P A G E | 37 III PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON 1 Introduction and Summary ............................................................................................. 39 2 Experimental Design ....................................................................................................... 41 2.1 Comparison of Label-Free and TMT-Based Quantification ..................................... 41 2.2 Proteoform-Directed Analysis .................................................................................. 42 3 Results .............................................................................................................................. 43 3.1 Evaluation of TMT Labeling Efficiency and Sample Consistency ........................... 43 3.2 Comparison of LFQ and TMT Proteomic Analyses ................................................. 45 3.3 Carbohydrate-Dependent Protein Abundance ........................................................ 47 3.4 Proteoform-Directed Top-Down Analysis ................................................................ 50 3.5 Analysis of Proteoform Termini ............................................................................... 51 3.6 Discovery-Based Open Modification Search ........................................................... 54 4 Discussion and Conclusion ............................................................................................ 57 4.1 Future Direction for Quantitative Analysis ............................................................... 57 4.2 Proteoform-Directed Analysis of B. thetaiotaomicron´s Proteome .......................... 59 III | PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON P A G E | 38 Parts of the following chapter have been published in “The intracellular proteome of the gut bacterium Bacteroides thetaiotaomicron is widely unaffected by a switch from glucose to sucrose as main carbohydrate source” Genth et al., Proteomics, 22(22), 1–6, (2022). Supplementary material for (Genth et al., 2022) Additional supplementary information’s is freely available for download at the publisher’s website https://www.doi.org/10.1002/pmic.202200189. The MS proteomics raw data and complete Proteome Discover search results have been deposited to the ProteomeXchange Consortium (http://www.proteomexchange.org/) via the PRIDE (Vizcaíno et al., 2014) partner repository with the data set identifier PXD033704. PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON | III P A G E | 39 1 Introduction and Summary The human gut microbiota significantly enhances the nutritional value of human diets by breaking down macromolecules and activating key intestinal genes to facilitate nutrient absorption (Ecklu-Mensah et al., 2022). Studies in germ-free mice have demonstrated that microbial colonization reduces the required caloric intake for weight maintenance by 30% and induces rapid changes in body fat composition (Wostmann et al., 1983; Bäckhed et al., 2004). Dietary choices can influence the composition of the gut microbiota and the immune response, both of which are key factors in the development and progression of inflammatory bowel disease (IBD) (Dolan and Chang, 2017). Individuals with IBD commonly develop increased sensitivity or intolerance to certain foods, leading them to avoid specific dietary components (Ballegaard et al., 1997; Zallot et al., 2013). Notably, the adoption of Western dietary habits, often associated with high levels of processed foods and refined carbohydrates, significantly increases the risk of gastrointestinal inflammatory disorders, including IBD (Ng, 2014; Khademi et al., 2021). However, it is important to acknowledge that dietary interventions can also be utilized to promote healthy gut microbial functions (Ecklu-Mensah et al., 2022). Within the human gut, Bacteroides represents one of the most abundant genera of bacteria (Arumugam et al., 2011). Bacteroides thetaiotaomicron, constituting approximately 6% of the total gut microbiota, plays a crucial role in reinforcing the mucosal barrier, maintaining immune response homeostasis, and processing nutrients (Zocco et al., 2007). Its complex repertoire of glycosylhydrolases enables the metabolism of a wide range of otherwise indigestible dietary polysaccharides and host-derived glycans in the human gut (Xu and Gordon, 2003). The impact of dietary sugars on the competitive dynamics within the human gut microbiota has great consequences. Prolonged consumption of a high-sugar diet can significantly alter gut microbial diversity (Do et al., 2018; Alasmar et al., 2023), leading to a displacement characterized by reduced levels of Bacteroidetes, similar to the dysbiosis observed in IBD (Frank et al., 2007). This dietary shift prompts bacteria to employ various adaptive mechanisms, which can include the application of specific carbohydrate transport and utilization systems or virulence genes that mediate toxin production or immune evasion (Poncet et al., 2009). While monosaccharides, such as glucose and fructose, may arrest the colonization of B. thetaiotaomicron in the gut by suppressing the expression of colonization-related proteins (Townsend et al., 2019), increased glucose intake in mice promotes mucolytic bacteria, including B. fragilis (Khan et al., 2020). This promotion could potentially compromise the integrity of the protective intestinal mucosal barrier, a critical factor in initiating intestinal inflammation (Png et al., 2010). The contradictory findings highlight the need for further research on this prominent genus and its sugar-microbiota interactions. III | PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON P A G E | 40 This chapter, prompted by the observed correlation between increased sugar consumption and the potential development of IBD (Khademi et al., 2021), aims to evaluate proteomic changes in B. thetaiotaomicron in response to different sugar sources. Aim of this study: - Utilize two quantitative bottom-up proteomics analyses – label-free quantification (LFQ) and isobaric labeling-based quantification using TMT – to evaluate proteomic changes induced by the presence of sucrose and glucose. - Perform a comparative analysis of quantitative results to determine the optimal methodology for future proteomic quantitative analysis. - Identify proteins whose abundance levels are influenced by the type of carbohydrate. - Apply various sample preparation techniques to deplete the high-molecular-weight proteome and increase the coverage of the low-molecular-weight proteome. - Conduct a proteoform-directed top-down analysis and employ a discovery-based open modification search to identify post-translational modifications and potential proteolytic cleavage events. PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON | III P A G E | 47 FIGURE III-7 | Analysis of the Proteomic Response between LFQ and TMT. (A) Comparison of log2 ratios of sucrose versus glucose for LFQ and TMT approaches. Proteins exhibiting higher abundance in the presence of either sucrose (32 proteins, red dots) or glucose (6 proteins, blue dots) are highlighted. The significance (p ≤ 0.05) was determined using a modified Pearson's correlation test (Yang et al., 2014). The red line represents the Pearson correlation (Pr), excluding one outlier. Direction pathway analysis of proteins of higher abundance in the presence of (B) sucrose or (C) glucose. Both graphs display the number of selected proteins in each COG category, the percentage of proteins in the specified category, and the corresponding p-values. 3.3 Carbohydrate-Dependent Protein Abundance Statistical analysis (two-sided Welch's t-test with Benjamini-Hochberg FDR correction for multiple testing, q ≤ 0.05 and Log2 fold change of ±0.485) identified 37 differentially abundant proteins for LFQ (FIGURE III-5A) and 32 for TMT (FIGURE III-5B). Visualization of the LFQ results revealed differential abundance of proteins associated with carbohydrate metabolism (red dots), amino acid transport and metabolism (yellow dots), and inorganic ion transport and -2 -1 0123456 -3 -2 -1 0 1 2 3 105 124 120 62 150 43 6 334 108 1.9E-26 8.3E-6 1.0E-5 7.5E-5 3.6E-4 1.1E-3 2.8E-3 2.1E-2 3.6E-2 134 83 29 334 15 75 64 48 2.0E-21 3.8E-5 1.2E-4 6.6E-4 1.2E-3 2.8E-3 3.4E-3 1.1E-2 TMT log2 (Sucrose/Glucose) LFQ log2 (Sucrose/Glucose) Pr = 0.81 B A C Carbohydrate transport / metabolism Amino acid transport / metabolism Energy production / conversion Inorganic ion transport / metabolism Cell wall/membrane/envelope biogenesis Lipid transport / metabolism Cell motility Function unknown Coenzyme transport / metabolism 020 40 60 80 100 Protein groups (%) p-value Translation, ribosomal, biogenesis Nucleotide transport / metabolism Signal transduction mechanisms Function unknown Secretion, vesicular transport Replication, recombination, repair Transcription PTM, protein turnover, chaperones 020 40 60 80 100 Protein groups (%) p-value III | PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON P A G E | 48 metabolism (blue dots) (FIGURE III-8A, B). Notably, certain proteins such as aldolase 1epimerase (Q8AAU2), pyruvate-phosphate dikinase (Q8AA21), fructokinase (Q8A6W9), glycine cleavage system H protein (Q8A4S8), and ROK family transcriptional repressor (Q8A4V4) consistently exhibited high intensity-based absolute quantification (iBAQ) values, regardless of the carbon source used in cultivation (FIGURE III-8C-D). These values, calculated by dividing a protein's total non-normalized intensity by the number of theoretically observable tryptic peptides for that protein, indicate the relative absolute abundance of proteins. These results suggest consistent abundance levels of certain proteins and may indicate their potential relevance in diverse cellular processes under different culture conditions. FIGURE III-8 | Overview of LFQ Quantitative Results. (A) Volcano plot with dashed vertical lines represents Log2 cutoffs and the dashed horizontal line represents a q-value of 0.05 (two-sided Welch's t-test with Benjamini-Hochberg FDR correction for multiple testing). (B) Proteins involved in amino acid, carbohydrate, or inorganic ion transport and metabolism according to COG annotations. iBAQ intensities of proteins identified in the presence of (C) glucose or (D) sucrose. Additionally, an increased abundance of proteins associated with fructan metabolism was identified, suggesting a specialized adaptation for the utilization of fructose-based oligosaccharides. This adaptation was facilitated by the starch utilization system (Sus), encoded within the polysaccharide utilization locus (PUL) 22, which includes eight open reading frames (FIGURE III-9B) (Martens et al., 2008; Sonnenburg et al., 2010). The regulation of this operon is mediated by a hybrid two-component system, positioned adjacent to the PUL, enhancing the efficiency of fructose-based carbohydrate utilization (Sonnenburg et al., 2010). 0500 1000 1500 2000 2 4 6 8 10 12 -2 0 2 4 6 8 0 1 2 3 4 5 6 7 0500 1000 1500 2000 2 4 6 8 10 12 12214 17 20 1 4 24 3 7 10 816 21 19 6 22 18 13 15 9 5 23 11 1 2 34 5 67 89 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 12 4 20 213 8 14 17 3 1 24 16 7 21 10 22 6 19 18 23 15 95 11 Glucose Log10 iBAQ Ranked proteins C -Log10 p-value LFQ Log2 (Glucose / Sucrose) A B Amino acid transport / metabolism 1) Aspartate aminotransferase 2) Glycine cleavage system H protein Carbohydrate transport / metabolism 3) Alpha-1,2-mannosidase 4) Fructokinase 5) Alpha-xylosidase 6) Glycoside hydrolase family 92 7) L-fucose isomerase 8) Levanase 9) Alpha-mannosidase 10) Beta-galactosidase 11) Alpha-1,2-mannosidase 12) ROK family transcriptional repressor 13) Glycoside hydrolase family 32 14) Aldose 1-epimerase 15) Beta-galactosidase 16) Alpha-1,2-mannosidase 17) Pyruvate-phosphate dikinase 18) Glucan 1,4-alpha-glucosidase SusB 19) 4-alpha-glucanotransferase Inorganic ion transport / metabolism 20) to 24) SusC homolog Sucrose Log10 iBAQ Ranked proteins D PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON | III P A G E | 49 Periplasmic hydrolysis of sucrose breaks it into its constituent monosaccharides, glucose, and fructose, which are then imported into the cytoplasm and directed toward central metabolic pathways (FIGURE III-9A). Glucose, a preferred energy source in numerous microorganisms, can access central metabolic pathways like glycolysis through the Embden-Meyerhof pathway. This pathway generates pyruvate, ATP, and precursor metabolites for the tricarboxylic acid (TCA) cycle. Unlike other carbohydrates, fructose could also directly enter the EmbdenMeyerhof pathway for energy generation, without the need for additional energy-consuming conversion steps (FIGURE III-9A). This metabolic pathway confers an advantage, enhancing the efficiency of fructose utilization across various bacteria, including B. thetaiotaomicron. FIGURE III-9 | Sucrose Utilization Pathway in B. thetaiotaomicron. (A) Sucrose uptake occurs via the Sus system and upon the cleavage of the O-glycosidic bond, glucose and fructose molecules are channeled into the major glycolytic pathways. The generated pyruvate enters the tricarboxylic acid (TCA) cycle for further processing. (B) Polysaccharide utilization locus (PUL) 22 encodes for the Sus system. Enzymes essential to this catabolic process are highlighted in gray and represented by the following abbreviations: HK (hexokinase), GPI (glucose-6-phosphate isomerase), FRK (fructokinase), and PFK (phosphofructokinase). O O O Sucrose Sus OO Glucose Fructose HK ATP ADP O P FRK ATP ADP O P O P GPI PFK 2 ATP 2 ADP Fructose-6-P Glucose6-P O P P O P PFructose1,6-bP Glycolysis TCA cycle Pyruvate PUL 22 1754 Hybrid two component system Fructokinase Monosaccharide import Glycoside hydrolases 32 SusE SusD homolog SusC homolog Glycoside hydrolase 32 1757 1758 1760 1762 17631761 17651759 A B III | PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON P A G E | 50 3.4 Proteoform-Directed Top-Down Analysis To increase the coverage of the low-molecularweight proteome, three distinct methods were employed. These methods included two HMWP depletion techniques, GELFrEE fractionation, and SCX fractionation. The number of identifications per method ranged from 360 to 540 protein groups (FIGURE III-10A) and 1,298 to 1,534 proteoforms (FIGURE III10B), respectively. Notably, all methods displayed comparable identification metrics, as demonstrated by the mean PrSMs per proteoform (FIGURE III-10D), residue cleavage (FIGURE III-10D), -Log (E-value) (FIGURE III10E), and the size distribution of the identified proteoforms (FIGURE III-10F). The overlap between the different methods ranged from 3% to 9%, and the majority of proteoforms were uniquely identified by individual methods, constituting approximately 25% each (FIGURE III-10G). Combining all three database searches resulted in the identification of 865 protein groups FIGURE III10A), represented by a total of 3,117 proteoforms (FIGURE III-10B). Detailed information on the exact identification metrics for each method is provided in Appendix 1.2. In summary, the number of identifications per SCX fraction ranged from 125 to 403 protein groups (FIGURE A-12A) and 185 to 1,280 proteoforms (FIGURE A-12B). The highest number of protein groups and proteoform identifications was achieved in fraction 3, which mostly contains higher charged species (FIGURE A-12A-B). Comparison of the acidic and basic HMWP depletion showed that comparable numbers of FIGURE III-10 | Top-down Methodology Comparison. Number of (A) isoforms, (B) proteoforms, (C) mean PrSMs per proteoform. Distribution of (D) residue cleavage, (E) –Log (E-value) and (F) theoretical proteoform mass. (G) Proteoform overlap between different methods. 360 409 540 865 25 33 19 34 1298 1321 1534 3117 HMWP depletion GELFrEE SCX Combined 0 200 400 600 800 1000 ∑ Isoforms A 0 10 20 30 40 Mean PrSMs per proteoform D 0 1000 2000 3000 4000 ∑ Proteoforms B C 0 20 40 60 80 100 Residue cleavage (%) 35 27 39 33 1.4 1.3 1.3 1.3 3.9 4.6 3.5 3.9 0.0 0.5 1.0 1.5 2.0 2.5 -Log10 (E-value) E 0 5 10 15 20 25 Theo. mass (kDa) F 710 (23%) 94 (3%) 779 (25%) 880 (28%) 288 (9%) 160 (5%) 206 (7%) SCX GELFrEE HMWP depletion G PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON | III P A G E | 51 protein groups and proteoform identifications were obtained (FIGURE A-12A-B), with each depletion method contributing a comparable number of unique identifications (FIGURE III-10AB). A comparison of GELFrEE fractions analyzed using different stationary phases (C4 and PLRP-S), indicated a generally higher number of protein groups (FIGURE A-12G) and proteoform identifications for the C4 column (FIGURE A-12H). The observed increase in the average number of PrSMs per proteoform identification of the PLRP-S column (FIGURE A-12I) can be attributed to decreased resolution due to wider elution profiles, resulting in peak broadening and poor separation of adjacent peaks compared to the C4 column (FIGURE A-12JK). These deviations are due to differences in column characteristics, including the larger particle size (5.0 µm, 1000 Å), shorter column length (17 cm), and potential dead volume introduced during crimping with ZIRCOFIT UHPLC fittings of the PLRP-S column, as opposed to the commercially purchased C4 column (2.6 µm, 150 Å, 50 cm). These differences may lead to fewer proteoforms being effectively separated in a given timeframe for the PLRP-S column. It is important to note that both columns operated at identical flow rates, gradients, and eluents (for detailed information, refer to chapter II.4.2). Therefore, optimization of gradient parameters, including the design of a nonlinear gradient for PLRP-S columns, provides an opportunity to improve chromatographic separation and increase the number of proteoforms (Trudgian et al., 2014). 3.5 Analysis of Proteoform Termini Out of 3,117 identified proteoforms, 2,942 (94%) had molecular weights between 2-10 kDa and consisted mostly of truncated proteoforms (FIGURE III-11A). N-terminal truncation accounted for 42% of the proteoforms, excluding N-terminal methionine excisions (NME), while C-terminal truncation accounted for 25% of the proteoforms (FIGURE III-11B). Notably, the heat shock protein (Q8AAA0) and the 60 kDa chaperonin (Q8A6P8) were prevalent in both TDP and BUP analyses, exhibiting numerous Nand C-terminal truncated proteoforms (38 and 36 respectively, TABLE A-2) and multiple peptides having high PSM counts (TABLE III-2). FIGURE III-11 | Neo-termini Analysis of Identified Proteoforms. (A) Size distribution and (B) percentage distribution of identified proteoforms, including Nand/or C-terminal truncated proteoforms with N-terminal initiator methionine excision and intact initiator methionine. 2 4 6 8 10 12 14 16 18 20 0 100 200 300 400 500 600 700 N-Term C-Term Intern Canonical ∑ Proteoforms Theo mass (kDa) C-term 25% Intern 28% N-term 42% Canonical 5% NME: Intact 65% 35% Excised A B III | PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON P A G E | 52 Internal proteoforms resulting from both Nand C-terminal truncation accounted for an additional 28%. Only 5% of proteoforms were classified as full-length canonical proteins, with 65% retaining and 35% lacking their N-terminal methionine (FIGURE III-11B). Approximately 15% of the proteoforms with alanine, glycine or serine positioned at the P1´ site following the initiator methionine retained the initiator methionine, resulting in 85% lacking the initiator methionine (FIGURE III-12A). Methionine cleavage was detected in all proteoforms exhibiting a proline at P1´ (FIGURE III-12B). Although lysine is the most common amino acid at P1´ (30.8%, FIGURE III-12C), and NME typically favors amino acids with smaller side chains at P1´(Frottin et al., 2006), a proteoform with lysine at P1´ was still exhibited NME cleavage (FIGURE III-12B). In total, 23 proteins were identified as both NME and non-NME proteoforms (FIGURE III-12D). Among them, the 60 kDa chaperonin (Q8A6P8) was identified with 15 proteoforms retaining the initial methionine and 13 lacking it (FIGURE III-12E). These proteoforms exhibited minor C-terminal amino acid truncations, possibly resulting from potential exopeptidase activity. This observation shows how a proteolytic event, such as NME, in conjunction with potential exopeptidase activity, can lead to the generation of multiple proteoforms (FIGURE III-12E). FIGURE III-12 | Analysis of Methionine Cleavage. (A) Number of identified proteoforms with N-terminal initiator methionine excisions (with NME) and intact initiator methionine (without NME). (B) Influence of the gyration radius of the P1´ amino acid residue on methionine cleavage. (C) Amino acid composition of the P1´ position following initiator methionine. (D) Overlap of proteins with and without the NME. (E) Identified proteoforms for 60 kDa chaperonin (Q8A6P8). G A S C T P V D N L I Q E H M F K Y W R 0 20 40 60 80 100 0% 10% 20% 30% G A S C T P V D N L I Q E H M F K Y W R 0 50 100 150 200 With NME Without NME Number of proteoforms Amino acid (P1´) A Met cleavage efficiency (%) Amino acid (P1′) 0.20 0.93 1.65 2.38 Side-chain gyration radius (Å) 0.00 C G I L V A P F H W Y C M K R D E N Q S T B 57 (26%) 23 (11%) 139 (63%) Without NME 162 proteins With NME 80 proteins [-].MA...KK.[G] [-].MA...GV.[D] [-].MA...AN.[A] [-].MA...AN.[A] [-].MA...NA.[V] [-].MA...VK.[V] [-].MA...KV.[T] [-].MA...VT.[L] [-].MA...TL.[G] [-].MA...LG.[P] [-].MA...GP.[K] [-].MA...PK.[G] [-].MA...KG.[R] [-].MA...GR.[N] [-].MA...AP.[H] [M].AK...VD.[A] [M].AK...VD.[A] [M].AK...DA.[L] [M].AK...AL.[A] [M].AK...LA.[N] [M].AK...AN.[A] [M].AK...NA.[V] [M].AK...AV.[K] [M].AK...KV.[T] [M].AK...LG.[P] [M].AK...RN.[V] [M].AK...KD.[G] [M].AK...QN.[T] [L].FN...VD.[A] [Y].FV...EA.[L] [V].RE...HA.[A] [V].RE...AA.[G] [V].RE...RV.[A] [N].AR...TE.[C] [C].VI...MM.[-] 0100 200 300 400 500 Sequence length Identified proteoforms Initiator methionine cleaved Initiator methionine not cleaved D E Amino acid (P1′) Number of proteoforms Amino acid (P1′) B A C Methionine cleavage (%) PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON | III P A G E | 53 Analysis of the N-terminal residues preceding the cleavage site (P1), and the C-terminal amino acids following it (P1´), revealed a uniform distribution of abundance for the majority of the detected neo-termini (FIGURE III-13). However, the detection of proteoforms with either aspartate at P1 or proline at P1´ led to the detection of a prominent Asp-Pro sequence logo in the acidic HMWP depletion, GELFrEE fractionation, and SCX fractionation experiments (FIGURE III-13 B-D). Conversely, in the HMWP deletion performed under basic conditions (pH 8.5), this specific sequence pattern was not observed (FIGURE III-13A). Instead, the cleavage data showed an increased abundance of two distinct proteoforms: one featuring asparagine at P1, and the other featuring either glycine or serine at P1´, resulting in noticeable sequence logos, Asn-Gly and Asn-Ser, respectively (FIGURE III-13A). The specific cleavage patterns, which varied depending on the method and pH of the depletion used, suggest that artificial cleavage events may have been introduced during sample preparation or LC-MS/MS analysis. FIGURE III-13 | Analysis of Proteoform Neo-Termini. Proteoform analysis for (A) basic HMWP depletion, (B) acidic HMWP depletion, (C) GELFrEE fractionation, and (D) SCX fractionation. Heatmaps illustrate the distribution and intensity of various proteoforms characterized by their N-terminal (P1) and C-terminal (P1´) residues, while bar plots summarize the total numbers of proteoforms identified. Count Count Amino acid (P1) Amino acid (P1' ) Count Count Amino acid (P1) Amino acid (P1' ) Count Count Amino acid (P1) Amino acid (P1' ) AB C D Count Count Amino acid (P1) Amino acid (P1' ) III | PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON P A G E | 54 3.6 Discovery-Based Open Modification Search During the preparation of biological samples and conducting LC-MS/MS analysis, artificial modifications, non-covalent adduct formation, and artificial truncation events may occur (Schaffer et al., 2021). To address this issue and to identify both known and novel modifications, a discovery-based open modification search was performed. This allowed for the identification of precursor mass shifts without a predefined list of PTMs. While several mass shifts hinted towards potential interesting PTMs on proteoforms of B. thetaiotaomicron, it's crucial to emphasize that these identifications are solely based on MS1 intensities, and mass shifts may also be attributed to inaccuracies in precursor mass deconvolution (Jeong et al., 2020). Therefore, the following results serve as preliminary indications and warrant further validation. Thousands of peptide spectrum matches (PrSMs) per mass shift were detected, each exhibiting varying degrees of mass accuracy (precursor mass tolerance: 10 ppm). This resulted in a broad distribution of delta precursor mass shifts. Following manual inspection of precursor isotope distribution and fragmentation spectra for potential a PTM, the monoisotopic mass corresponding to entries in the PSI-MOD (MontecchiPalazzi et al., 2008) or UniMod database (Creasy and Cottrell, 2004) will be reported. The detected mass shifts included methionine oxidation (+15.994 Da), cysteine dioxidation (+31.988 Da), and combinations of multiple oxidations (FIGURE III-14). These modifications can occur spontaneously during sample preparation and storage (Kaulich et al., 2022b). Additionally, certain mass shifts indicate the absence of specific amino acid residues, such as cysteinyl (+103.009 Da), phenylalaninyl (+147.068 Da), and tyrosinyl (+163.063 Da). Frequent misassignments were often identified as either incorrect annotation (-131.040 Da) or absence (+131.040 Da) of N-terminal initiator methionine residues. Incorrect mass shifts, arising from multiple PTMs or incorrectly assigned modifications, such as the absence of initiator methionine and misassignments of acetylation (-42.010 Da), contributed to prevalent mass shifts of +89.03 Da (FIGURE III-14 and FIGURE A-13A). FIGURE III-14 | Distribution of Precursor Delta Mass Shifts. Arrows highlight potential PTMs with matched monoisotopic masses according to the PSI-MOD (Montecchi-Palazzi et al., 2008) or UniMod (Creasy and Cottrell, 2004) database. -200 -150 -100 -50 050 100 150 200 0 2500 5000 7500 10000 60000 PrSMs Δ Precursor mass shift (Da) Methyl- (14 Da) Oxidation (16 Da) Methionyl- (131 Da) Acetyl (-42 Da) (89 Da) Methionyl- (131 Da) Tyrosinyl- (163 Da) Phenylalanyl- (147 Da) Cysteinyl- (103 Da) Dioxidation (32 Da) Formyl- (-28 Da) Acetyl- (-42 Da) Methionyl- (-131 Da) Trioxidation (48 Da) PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON | III P A G E | 55 The formation of disulfide bonds in cysteine-containing proteoforms can be indicated by small mass shifts, such as –2 Da and –4 Da. Several potential disulfide bonds have been identified in proteins such as 50S ribosomal protein L32 (Q8A138), which displays a zinc finger motif (Cys-Xaa2-Cys-Xaa9-Cys-Xaa2-Cys) (FIGURE A-13B), and peptidyl-prolyl cis-trans isomerase (Q8A607), featuring an eight-cysteine motif (8CM) (FIGURE A-13C). Although 8CM motifs are present in a large number of fungal extracellular membrane proteins (Kulkarni et al., 2003) and plant defensins (José-Estanyol et al., 2004), their role in bacterial proteins is unclear. A mass shift of 45.987 Da was detected on ribosomal protein S12 (Q8A472), suggesting a potential β-methylthio-aspartic acid modification (FIGURE III-15A). This PTM has been reported in other bacterial ribosomal protein S12 proteins (Kowalak and Walsh, 1996; Strader et al., 2004, 2011). Furthermore, the potential modification was detected in a BUP search, on the same aspartic acid within a peptide spanning the same sequence (FIGURE A-13D). Proteoforms of the glycine cleavage system H protein (Q8A4S8) exhibited a mass shift of 188.033 Da, potentially indicating an N6-lipoyllysine modification (FIGURE III-15B). This PTM is essential for transferring a methylamine group from the P protein to the T protein in the glycine cleavage system, generating CO2, NH3, and N5,N10-methylene-tetrahydrofolate (THF) (McCarthy and Booker, 2020). FIGURE III-15 | Putative PTMs in B. thetaiotaomicron proteins. (A) β-methylthio-aspartic acid on ribosomal protein S12 (Q8A472). (B) N6-lipoyllysine on glycine cleavage system H protein (Q8A4S8). (C) O-(pantetheine 4'-phosphoryl)serine on acyl carrier protein (Q8A2E6). Distal thiol Ribosomal protein S12 β-methylthio-aspartic acid (Mass Difference: 0.005 Da & 0.34 ppm, P-Score: 2.3E-58, Residue cleavage: 23%) Glycine cleavage system H protein N6-lipoyllysine (Mass Difference: 0.058 Da & 3.96 ppm, P-Score: 1.0E-35, Residue cleavage: 16%) Acyl carrier protein O-(pantetheine 4'-phosphoryl)serine (Mass Difference: 0.005 Da & 0.56 ppm, P-Score: 1.3E-90, Residue cleavage: 61%) NPTIQQLVRKGREVLVEKSKSPALD 25 S 26 CPQRRGVCVRVYTTTPKKPNSAMR 50 K 51 VARVRLTNQKEVNSYIPGEGHNLQ 75 E 76 HSIVLVRGGRVK LPGVRYHIVRG 100 T 101 LDTAGVAGRTQRRSKYGAKRPKPG 125 Q 126 AAPAKKKC D NMNFPQNLKYTNEHEWIRVEGDIAY 25 V 26 GITDYAQEQLGDIVFVDIPTVGET 50 L 51 EAGETFGTIEVV TISDLFLPLAG 75 E 76 ILEQNEALEENPELVNKDPYGEGW 100 L 101 IKMKPADASAAEDLLDAEAYKAVV 125 N 126 GC K NSEIASRVKAIIVDKLGVEESEVTN 25 E 26 ASFTNDLGAD LDTVELIMEFEKE 50 F 51 GISIPDDQAEKIGTVGDAVSYIEE 75 H 76 A K C S HN ab S g OOH O CH3 H N O S S O HN HN OP O OH O CH3 H3C OH HN O N H HS O O A B C III | PROTEOMIC ANALYSIS OF B. THETAIOTAOMICRON P A G E | 56 The largest and most abundant potential PTM discovered was O-(pantetheine 4'- phosphoryl)serine (340,085 Da) on the acyl carrier protein (ACP) (Q8A2E6) (FIGURE III-15C). The prosthetic moiety, 4'-phosphopantetheinyl (Ppant), has a highly reactive distal thiol group that facilitates the covalent transport of various chemical groups, including acyl groups, during fatty acid biosynthesis (Chan and Vogel, 2010). Interestingly, ACP proteoforms exhibiting the Ppant modification showed a wide range of mass shifts that may be caused by different PTMs (FIGURE III-16). Some of these mass shifts may be attributed to common acyl intermediates during fatty acid biosynthesis, such as acetyl (+42,010 Da) and malonyl (+86,000 Da) (Chan and Vogel, 2010). Since fatty acid biosynthesis involves multiple steps to produce full-length chains (typically C16 or C18) (Cronan and Thomas, 2009), the presence of larger mass shifts may indicate longer chain-length intermediates. Furthermore, several mass shifts may indicate thiol-related modifications, such as cysteinylation (119.004 Da) and sulfur dioxide (SO2) addition (+63.961 Da) (FIGURE III-16). FIGURE III-16 | Distribution of Precursor Mass Shifts on Acyl Carrier Protein (Q8A2E6). All indicated mass shifts additionally exhibit the O-(pantetheine 4'-phosphoryl)serine mass shift of 340.085 Da. Arrows highlight potential PTMs with matched monoisotopic masses according to PSI-MOD (MontecchiPalazzi et al., 2008) or UniMod (Creasy and Cottrell, 2004) databases. 050 100 150 200 250 300 0 500 1000 2500 PrSMs Δ Precursor mass shift (Da) 16 Da SOH 42 Da SCH3 O 64 Da SSH OO S O OH O 86 Da 119 Da SS NH2 OH O O O NH P O OH O H3CCH3 HO NHO N H SH O Potential intermediates of fatty acid synthesis 4’-Phosphopantetheinyl P A G E | 63 IV INFLUENCE OF PH ON BACTERIAL PROTEOMES 1 Introduction and Summary ............................................................................................. 64 2 Experimental Design ....................................................................................................... 66 2.1 Bottom-up LFQ Analysis of HGM Proteomes .......................................................... 66 2.2 Top-down LFQ Analysis of B. producta ................................................................... 66 3 Results .............................................................................................................................. 68 3.1 Proteomic Adaptations to Acidic and Alkaline pH in B. longum .............................. 68 3.2 Analysis of the Acidic Response of three HGM Members ....................................... 71 3.3 Comparison of LMWP-Top-Down and Full-Proteome Bottom-Up Analysis for B. producta ........................................................................................................ 79 3.4 Potential pH-induced Asp-Pro Cleavage ................................................................. 84 3.5 Phosphorylation of HPr proteins .............................................................................. 87 4 Discussion and Conclusion ............................................................................................ 89 4.1 Quantitative analysis of the B. longum proteome .................................................... 89 4.2 Quantitative analysis of the B. thetaiotaomicron proteome ..................................... 92 4.3 Quantitative analysis of the B. producta proteome .................................................. 93 4.4 Asp-Pro peptide bond hydrolysis ............................................................................. 98 IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 64 1 Introduction and Summary The human gut microbiome (HGM), comprising diverse microbial communities, plays a crucial role in maintaining health. Its composition and balance are influenced by various factors, including growth factors, micronutrients, antimicrobial compounds, and gut pH (Rodionov et al., 2019; Beam et al., 2021; Firrman et al., 2022). While some variability in the pH of the gastrointestinal tract is normal due to factors like diet and microbial activity, significant deviations from typical ranges can alter the microbiome’s composition and function, potentially impacting health (Firrman et al., 2022). Generally, the proximal small bowel has lower pH levels (pH 5.5 to pH 7.0) compared to the descending and rectosigmoid colon, which maintains slightly higher pH levels (pH 6.6 to pH 7.5) (Nugent et al., 2001). Notably, increased colonic acidity has been associated with various gastrointestinal diseases, including irritable bowel syndrome (IBS) and inflammatory bowel disease (IBD) (Nugent et al., 2001; Ringel-Kulka et al., 2015). Microbial responses to varying pH conditions involve physiological and molecular adaptations aimed at maintaining intracellular pH (pHi) homeostasis. These adaptations can include proton translocation by specialized pumps such as ATP synthase and small ion (Na⁺, K⁺, or Ca²⁺)/H⁺ antiporters (Krulwich et al., 2011). Additionally, enzyme-catalyzed reactions, including processes such as decarboxylation, consume protons, whereas deaminase-catalyzed reactions increase the concentration of alkaline compounds (Krulwich et al., 2011). Protective mechanisms also include changes in lipid composition to reduce proton permeability of the cell, promotion of biofilm formation, adjustment of cell density, and implementation of repair mechanisms to counteract increased damage to macromolecules (Guan and Liu, 2020). Despite extensive investigations into microbial acid stress responses, knowledge regarding the proteomic adaptations of specific HGM members to acidic stress remains limited. Understanding these proteomic changes could provide valuable insights into how these microbes maintain resilience and adapt to acidic environments. Addressing this research gap could advance understanding of gut microbial dynamics and their potential implications for human health, as well as facilitate the development of targeted therapeutic strategies for gastrointestinal diseases. INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 65 Aim of this study: - To identify proteomic alterations in response to diverse pH conditions (pH 6.0, pH 7.0, and pH 8.0), label-free quantification analyses using bottom-up proteomics were conducted on three HGM species: Bacteroides thetaiotaomicron, Blautia producta, and Bifidobacterium longum. - Compare protein abundances to identify proteins exhibiting co-abundance in both acidic and alkaline responses, as well as those exhibiting specific abundance under acidic or alkaline pH conditions relative to growth at pH 7.0. - Identify concurrent and distinct proteomic alterations in pathway-related or stress-associated proteins among the three bacterial species. - Utilize a top-down proteomics label-free quantification analysis to quantify proteoforms. - Perform a comparative analysis of bottom-up and top-down quantitative results to evaluate their capacity to quantify the same proteins, whether differentially abundant or not. - Conduct a discovery-based open modification search to identify potential post-translational modifications. IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 66 2 Experimental Design 2.1 Bottom-up LFQ Analysis of HGM Proteomes To analyze the proteomic response to cultivation at different pH values in B. thetaiotaomicron, B. producta, and B. longum, the bacteria were cultured in YCFA medium at acidic (pH 6.0) and neutral (pH 7.0) conditions in five biological replicates. B. longum was additionally cultivated under alkaline (pH 8.0) conditions (FIGURE IV-1, chapter II.2.1). At mid-stationary phase, cells were harvested and lysed using freeze-thawing and proteins were cleaned up using ethanol precipitation (chapter II.3.1). Intracellular proteomes were isolated, digested using trypsin with a 1:40 enzyme-to-substrate ratio and subjected to solid-phase extraction (chapter II.3.4). All samples were separated online by reversed-phase chromatography with a gradient of 120 minutes and measured in triplicate on the Q-Exactive Plus mass spectrometer (chapter II.4.1). The acquired raw data were searched against the respective reference proteomes using PD 2.5 (chapter II.5.1). Median normalization was used to center the data distribution and compensate for general concentration differences that could be caused by minor variations in sample concentration (chapter II. 5.4). The raw and normalized intensity data were tested for normal distribution and Pearson correlation (chapter II. 5.4). Statistical analysis (two-sided Welch's t-test with Benjamini-Hochberg FDR correction for multiple testing, q ≤ 0.05) was performed on high-confidence protein identifications (1% FDR) with at least three quantitative values out of the five biological replicates. FIGURE IV-1 | Experimental Design to Analyze the pH-Induced Proteomic Response of HGM Bacteria. Bottom-up proteomic LFQ analysis of B. thetaiotaomicron, B. producta, and B. longum cultivated in YCFA medium at pH 6.0 and pH 7.0, with additional cultivation of B. longum at pH 8.0. 2.2 Top-down LFQ Analysis of B. producta To perform label-free quantification of proteoforms in B. producta, three of the five biological replicates grown in YCFA medium under different pH conditions (pH 6.0 and pH 7.0) were processed using acidic and basic depletion of the high molecular-weight proteome (FIGURE IV2) (Cassidy et al., 2019). This approach aimed to enhance the identification of low-molecularweight proteoforms. Details about the sample preparation are described in chapter II.3.2. Samples were separated by reversed-phase chromatography using a 60 min gradient and 3x LC-MS/MS 120 min Intensity-based quan. MS1 PD 2.5 Data analysisHGM member lysis Digestion SPE pH 8.0 pH 6.0 pH 7.0 Only B. longum INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 67 analyzed in duplicates on the Fusion Lumos mass spectrometer, employing the multi-CV FAIMS method which four distinct CVs (-60, -50, -40 and -20, chapter II.4.2) (Kaulich et al., 2022a). The acquired raw data were filtered using Freestyle v.1.6 based on the applied CVs, dividing each LC-FAIMS-MS/MS analysis into four distinct raw files. Subsequently, these files were searched against the B. producta reference proteome using PD 2.5 (chapter II.5.1). Quantification was performed utilizing the high-resolution feature detector node with the sliding window deconvolution algorithm with an average retention time width of 0.33 min. All results were subjected to an FDR correction for both PrSMs and proteoforms, with a threshold of 1%. Identified proteoforms were required to have a minimum C-Score of 40 (Leduc et al., 2014). Top-down data processing steps included CV summation, median calculation, and median normalization. The raw and normalized intensity data were tested for normal distribution and Pearson correlation (chapter II. 5.4). Statistical analysis (two-sided Student's t-test with permutation-based FDR correction for multiple testing, q ≤ 0.05) required at least two quantitative values out of three biological replicates. FIGURE IV-2 | Experimental Design for Top-down LFQ Analysis of the B. producta Proteome. B. producta cultivated in YCFA medium at pH 6.0 and pH 7.0 were subjected to acidic and basic HMWP depletion for top-down label-free quantification. 2x LC-MS/MS 60 min Intensity-based quan. MS1 PD 2.5 Data analysisB. producta lysis HMWP depletion pH 6.0 pH 7.0 Multi CV FAIMS IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 68 3 Results 3.1 Proteomic Adaptations to Acidic and Alkaline pH in B. longum Data Processing – Proteomic changes of B. longum after cultivation at pH 6.0, pH 7.0, and pH 8.0 were analyzed by BUP-LFQ analysis. A total of 991 proteins, covering 57% of the encoded B. longum proteome, were identified. Protein abundance profiles for both biological and technical replicates (FIGURE A-14), as well as culture pH comparisons (pH 7.0/pH 6.0, pH 7.0/pH 8.0, and pH 8.0/pH 6.0), were normalized to a median value of zero (FIGURE IV-3AB). The datasets showed low variability, with median coefficients of variation ranging from 3.8% to 6.7% (FIGURE IV-3C). Strong concordance among biological and technical replicates is demonstrated by high average Pearson correlation coefficients of 0.95 (FIGURE A-14E). Principal-component analysis (PCA) revealed pH-dependent separation and clustering of biological replicates, with the first two principal components accounting for 78% of the variation (FIGURE IV-3D). FIGURE IV-3 | Evaluation of Data Processing of the B. longum Proteome. Distribution of Log2 ratios of (A) raw data and (B) median normalized data. Each histogram is overlaid with a Gaussian distribution curve. The median (M) of the datasets is displayed in the upper left corner of each histogram. (C) Box-and-whisker plots of the coefficient of variation of the identified proteomes. Boxes capture the lower and upper quartiles with the median displayed as a horizontal line in the middle; whiskers represent the 1–99 percentile. (D) Principal component (PC) analysis with each circle represents a biological replicate cultivated at the indicated pH. Quantitative data – This study aimed to examine the proteomic response of B. longum to acidic (pH 6.0) and alkaline (pH 8.0) culture conditions by comparing protein abundance ratios pH 7.0 vs. pH 6.0 (indicating the acidic response) and pH 7.0 vs. pH 8.0 (indicating the alkaline response). A total of 933 and 935 proteins were quantified for the acidic and alkaline responses, respectively. Subsequent statistical analysis, employing a two-sided Welch's t-test -6 -4 -2 0 2 4 6 0 100 200 300 pH 7.0/pH 6.0 ∑ Proteins Log2 ratio M= -0.19 M= 0.00 M= -0.22 -6 -4 -2 0 2 4 6 0 100 200 300 pH 7.0/pH 8.0 ∑ Proteins -6 -4 -2 0 2 4 6 0 100 200 300 pH 8.0/pH 6.0 ∑ Proteins -6 -4 -2 0 2 4 6 0 100 200 300 pH 7.0/pH 6.0 ∑ Proteins Log2 ratio M= -0.46 M= 1.36 M= -1.85 -6 -4 -2 0 2 4 6 0 100 200 300 pH 7.0/pH 8.0 ∑ Proteins -6 -4 -2 0 2 4 6 0 100 200 300 pH 8.0/pH 6.0 ∑ Proteins BA C pH 6.0 pH 7.0 pH 8.0 0 5 10 15 20 60 80 Coefficients of variation (%) D -10 010 20 -10 0 10 20 pH 6.0 pH 7.0 pH 8.0 PC 2 (32% variance) PC 1 (46% variance) INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 69 with Benjamini-Hochberg FDR correction (FDR ≤ 0.05) and a Log2 fold change threshold of ±0.485, identified 444 differentially abundant proteins for the acidic response and 332 for the alkaline response. Changes in protein abundance were detected for proteins associated with cellular processes such as protein export, folding, degradation, DNA protection, and repair, as well as proteins of the translational machinery such as ribosomal subunits and aminoacyl-tRNA synthetases (FIGURE IV-4A-B). Additionally, several proteins involved in amino acid metabolism, fatty acid synthesis, and peptidoglycan biosynthesis were detected as differentially abundant (FIGURE IV-4C-D). The acidic proteomic response will be later described and discussed (see chapter IV.3.2). Therefore, the following results focus on the alkaline response and functional classes of proteins that had differential abundances in both the acidic and alkaline response and those that had differential abundance at ether acid or alkaline pH compared to growth at pH 7.0. At pH 8.0, a higher abundance of proteins involved in methionine and cysteine biosynthesis or interconversion was observed, including homoserine O-acetyltransferase (MetA; Q8G7A5), cystathionine γ-synthase (MetB; Q8G565), cystathionine β-synthase (Cbs; Q8G564), and methionine synthase (MetE; Q8G651). FIGURE IV-4 | B. longum Acidic and Alkaline Proteomic Response. (A-B) Volcano plots of the acidic (pH 7.0/pH 6.0) and (C-D) alkaline (pH 7.0/pH 8.0) response. Proteins are labeled by their respective gene names and are color-coded according to their involvement in cellular processes, with "PUF" denoting proteins of unknown function. Dashed vertical lines represent Log2 cutoffs, while the dashed horizontal line corresponds to a q-value of 0.05 (Two-sided Welch's t-test, corrected for multiple testing by Benjamini-Hochberg FDR calculation). q=0.05 D A q=0.05 Aminoacyl tRNA synthetases Protein folding and degradation Methyltransferases Protection and repair of DNA Ribosomal subunits Miscellaneous Cys & Met metabolism Fatty acid biosynthesis BCAA, Gly, Ser and Thr metabolism His, Phe, Tyr and Trp metabolism Pro biosynthesis Peptidoglycan biosynthesis B q=0.05 q=0.05 C ClpB GroS GrpE UvrD XseA XseB PtsH SigH PUF PUF PUF Cah Impa 0 1 2 3 4 5 6 7 8 9 -6 -4 -2 0 2 4 6 -Log p-value Log2(pH 7.0/pH 6.0) PUFs Protein export ThrS TrmB LepB LigA RpmF RpsL PUF Lacl PUF PUFs GcvH Cah PUF Impa 0 1 2 3 4 5 6 7 8 9 -6 -4 -2 0 2 4 6 -Log10 p-value Log2(pH 7.0/pH 8.0) MetA Cbs HutH HisI HisH HisE AroA AroB FabG 0 1 2 3 4 5 6 7 8 9 -6 -4 -2 0246 -Log10 p-value Log2ratio (pH 7.0/pH 6.0) Cbs MetB MetA MiaB TrmB MiaA FabG FabG Z AroA AroE AroG TyrA Z Z HutH HisH 0 1 2 3 4 5 6 7 8 9 -6 -4 -2 0 2 4 6 -Log10 p-value Log2(pH 7.0/pH 8.0) IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 70 Additionally, two S-adenosyl-L-methionine-dependent enzymes, tRNA-methyltransferases (TrmB; Q8G3T4) and tRNA-methylthiotransferase (MiaB; Q8G4H4), were more abundant at pH 8.0. At pH 7.0, proteins of higher abundance included ribosomal subunits, proteins of unknown function (PUFs; Q8G7V3, Q8G449, Q8G6I8, and Q8G6N5), and histidine synthesis enzymes such as imidazole glycerol phosphate synthase (HisH; Q8G4S6), phosphoribosylATP pyrophosphatase (HisE; Q8G694), and phosphoribosyl-AMP cyclohydrolase (HisI; Q8G6F6) (FIGURE IV-4). Conversely, histidine ammonia-lyase (HutH), which is involved in histidine degradation, and several aminoacyl-tRNA synthetases exhibited higher abundance at both pH 6.0 and pH 8.0 (FIGURE IV-4). By directional analysis (Yang et al., 2014) of the 932 shared quantified proteins between the acidic and alkaline responses significant variations in protein abundance and pathway changes across different pH environments could be identified. Differential changes in protein abundance (p ≤ 0.05) were categorized into four groups (FIGURE IV-5A). FIGURE IV-5 | Comparison of the Acidic and Alkaline Response of the B. longum Proteome. (A) Directional analysis with differentially abundant proteins colored based on their change in abundance (p ≤ 0.05), using a modified Pearson’s correlation test (Yang et al., 2014). (B) Directional pathway analysis on proteins with shared abundance at pH 7.0 or (C) shared abundance at pH 6.0 and pH 8.0 using COG annotations. The number and percentage of proteins and the corresponding p-values of each category are shown. -6 -4 -2 0 2 4 6 -4 -2 0 2 4 6 125 110 72 48 23 8.6E-11 1.7E-2 1.7E-2 2.6E-2 5.0E-2 68 119 80 45 51 9.0E-9 4.3E-4 6.1E-4 7.4E-4 2.3E-2 Translation, ribosomal structure and biogenesis Function unknown Transcription Cell wall/membrane/envelope biogenesis Defense mechanisms 020 40 60 80 100 Protein groups (%) p-value Nucleotide transport and metabolism Amino acid transport and metabolism Carbohydrate transport and metabolism Energy production and conversion Coenzyme transport and metabolism 020 40 60 80 100 Protein groups (%) p-value pH specific response Opposite response Shared respsone pH 7.0 Shared respsone pH 6.0 & pH 8.0 B Log2 (pH 7.0/pH 6.0) Log2 (pH 7.0/pH 8.0) C A INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 71 Firstly, 65 proteins (30%) had pH-specific responses, meaning they had differential changes in abundance only at a specific pH condition, either pH 6.0 or pH 8.0. Second, 14 proteins (6%) exhibited unspecific responses, increasing their abundance in one condition but decreasing in another. Examples of these proteins include cystathionine β-synthase (Cbs; Q8G564), nitrogen regulatory protein N-II (GlnB; Q8G738), and 30S ribosomal protein S12 (RpsL; P59162), which increased in abundance under acidic conditions and decreased under alkaline conditions. Third, 71 proteins (32%) maintained a consistent increase in abundance specifically at pH 7.0. Lastly, another 71 proteins (32%) exhibited a shared response at both pH 6.0 and pH 8.0, maintaining a consistent increase in abundance under both acidic and alkaline conditions. The Pearson correlation coefficient between the acidic and alkaline response of 0.59 suggests a moderately complementary proteomic response to acidic and alkaline conditions (FIGURE IV-5A). Directional pathway analysis on proteins with consistent increase in abundance specifically at pH 7.0 revealed significant enrichment (p < 0.01) in translational and transcriptional processes, as well as cell wall biogenesis, defense mechanisms, and proteins of unknown function (FIGURE IV-5B). Proteins with a consistent increase in abundance at pH 6.0 and pH 8.0 were enriched in COG categories involved in the transport and metabolism of nucleotides, amino acids, carbohydrates, and coenzymes, as well as proteins associated with energy production and conversion (FIGURE IV-5C). These include 5-enolpyruvylshikimate-3-phosphate (EPSP) synthase (AroA; Q8G5N6) (TABLE A-3), which is essential for catalyzing the formation of EPSP and inorganic phosphate from shikimate-3phosphate and phosphoenolpyruvate in the shikimate pathway. 3.2 Analysis of the Acidic Response of three HGM Members Data Processing – A total of 1,497, 1,772, and 927 proteins were identified in the B. thetaiotaomicron, B. producta, and B. longum datasets, respectively (FIGURE IV-6A-C). An overview of the datasets before and after median normalization and the Pearson correlation analysis is provided in the Appendix (FIGURE A-15). Proteins quantified in three out of five replicates of at least one culture condition accounted for 30% (1449 out of 4782) of all encoded proteins for B. thetaiotaomicron, 31% (1663 of 5365) for B. producta, and 52% (889 of 1725) for B. longum (FIGURE IV-6A-C). FIGURE IV-6 | Protein Identification Overview. The total number of proteins encoded in the genome, identified proteins, identified proteins, quantified proteins (detected in three of five replicates) and differentially abundant proteins for (A) B. thetaiotaomicron, (B) B producta and (C) B longum. 528 1449 1737 4782 Diff. abundant Quantified Identified Total proteins 020 40 60 80 100 Proteins filtered (%) A 502 1663 1772 5365 Diff. abundant Quantified Identified Total proteins 020 40 60 80 100 Proteins filtered (%) B C 446 889 927 1725 Diff. abundant Quantified Identified Total proteins 020 40 60 80 100 Proteins filtered (%) IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 72 Statistical analysis of the acid pH response (pH 7.0 vs. pH 6.0) of the three bacteria identified 528, 501, and 446 proteins as differentially abundant for B. thetaiotaomicron, B. producta, and B. longum, respectively (FIGURE IV-6A-C). Quantitative Data – Significant proteomic adaptations to acidic culture conditions were observed across all HGM bacteria (FIGURE IV-7A-C). Although a FDR of 5% was chosen as the threshold for statistical significance, the majority of proteins with differential abundance (Log2 fold change threshold of ±0.485) were identified with an FDR of 1% (FIGURE IV-7D). In B. thetaiotaomicron, an equal number of differentially abundant proteins were detected with increased abundance at both pH 6.0 and pH 7.0 (FIGURE IV-7D). In contrast, B. producta and B. longum exhibited a higher number of proteins with increased abundance at pH 6.0 (FIGURE IV-7D). Most significant changes involved proteins involved in metabolic pathways such as carbohydrate utilization, amino acid biosynthesis and degradation, purine and pyrimidine biosynthesis, and peptidoglycan synthesis. Additionally, proteins involved in cellular respiration, transcriptional and translational processes, and the protection and repair of macromolecules exhibited variations in abundance. A summary of differentially abundant proteins and their respective fold changes in the three HGM bacteria is provided in TABLE A-4, and explained in detail in the following sections. FIGURE IV-7 | Quantitative Proteome Analysis of Human Gut Bacteria. Volcano plot for all quantified proteins of (A) B. thetaiotaomicron, (B) B. producta, and (C) B. longum at pH 7.0 vs. pH 6.0. The dashed vertical lines represent Log2 ratio thresholds and the dashed horizontal lines represent a q-value of 0.05 or 0.01 (Two-sided Welch's t-test, Benjamini-Hochberg FDR corrected). (D) The total number of differentially abundant proteins quantified with increased abundance at pH 7.0 and pH 6.0 with a q-value of 0.05 or 0.01. 0 1 2 3 4 5 6 7 8 -6 -4 -2 0 2 4 6 -Log10 p-value Log2ratio (pH 7.0/pH 6.0) 0 1 2 3 4 5 6 7 8 -6 -4 -2 0 2 4 6 -Log10 p-value Log2ratio (pH 7.0/pH 6.0) 0 1 2 3 4 5 6 7 8 -6 -4 -2 0 2 4 6 -Log10 p-value Log2ratio (pH 7.0/pH 6.0) q=0.05 D q=0.01 q=0.05 q=0.01 q=0.05 q=0.01 A CB 273 201 191 255 301 255 -400 -200 0 200 400 B. thetaiotaomicron B. producta B. longum Number of differentially abundant proteins Increased pH 7.0 (5% FDR) Increased pH 7.0 (1% FDR) Increased pH 6.0 (1% FDR) Increased pH 6.0 (5% FDR) pH 6.0 pH 7.0 pH 6.0 pH 7.0 pH 6.0 pH 7.0 INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 79 had an increased abundance of the cell shape-determining protein (MreB; A0A2S4GRY8), capsule biosynthesis protein (CapA; A0A7G5MP92) (TABLE A-4), and several stage 0 and III sporulation proteins at pH 6.0 (TABLE A-5). CapA facilitates the addition of cell surface capsular polysaccharide that can protect bacteria from immune responses and alter host physiology (Porter and Martens, 2017). While the teichoic acid biosynthesis protein (A0A7G5MS69) was more abundant at pH 7.0, the D-alanyl-lipoteichoic acid biosynthesis protein (DltD; A0A7G5MQG2) increased ninefold at pH 6.0 (TABLE A-4). Furthermore, several Fts proteins involved in cell elongation, cell cycle control, and cell division were more abundant at pH 6.0 (TABLE A-4). Notably, the cell division protein FtsL (A0A4P6M6W0) showed the highest change in abundance in the dataset, with a 16-fold increase at pH 6.0. Acidic conditions have been reported to reduce the cell length of E. coli by modulating the division machinery, specifically the terminal cell division protein FtsN (Mueller et al., 2020). Although these findings originate from a Gram-negative bacterium and may not directly apply to B. producta, a Gram-positive bacterium lacking genomic data for FtsN, studies on other Gram-positive bacteria like S. aureus and S. pneumoniae, which also lack identifiable FtsN homologs, have demonstrated significant size alterations in response to pH variations (Perez et al., 2019; Mueller et al., 2020). These observations suggest the hypothesis that the cellular morphology of B. producta might also change in response to pH variations. However, initial electron microscope experiments conducted by Kathrin Schäfer (Department of Infectious Diseases and Microbiology, University of Lübeck, UKSH Lübeck; chaired by Prof. Dr. Jan Rupp) on B. producta in response to various culture pH values did not reveal significant changes in cell morphology. Although these preliminary findings do not conclusively rule out the possibility that the cellular morphology of B. producta changes under these specific conditions, further research is required to explore this hypothesis. 3.3 Comparison of LMWP-Top-Down and Full-Proteome Bottom-Up Analysis for B. producta Data Processing – Top-down proteomic analysis of acidic and basic HMWP depletions of B. producta identified 923 and 818 proteoforms (FIGURE IV-11A), corresponding to 211 and 190 protein groups (FIGURE IV-11B), respectively. The majority of identified proteoforms were between 2-10 kDa in molecular weight and were truncated versions of the canonical proteins (FIGURE IV-11C and D). N-terminal truncation accounted for 31% ± 4% of proteoforms, excluding N-terminal methionine excisions (NME) (FIGURE IV-11B and D). Additionally, Cterminal truncation represented 28% ± 3%, while internal proteoforms resulting from both Nand C-terminal truncation contributed an additional 30 ± 8%. This left 12% of identified IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 80 proteoforms as full-length canonical proteins, with 66% ± 3% retaining their N-terminal methionine and 34% ± 3% lacking it (FIGURE IV-11E and F). Interestingly, only 404 and 356 proteoforms had at least one quantitation value after acidic and basic HMWP depletion, respectively (FIGURE IV-11A). Both datasets had a high percentage of missing values, with 38% and 42% missing values after acidic and basic HMWP depletion, respectively. After data processing steps including CV summation, median normalization, and filtering for at least two quantitative values per pH condition, the datasets were reduced to 175 proteoforms after acidic and 135 after basic depletion, leaving many proteoforms unquantified (FIGURE IV-11A). FIGURE IV-11 | Identified Proteoforms using the Acidic and Basic HMWP Depletion Method. Number of identified, quantified, and differentially abundant (A) proteoforms and (B) protein groups. Distribution of identified neo-termini by proteoform size for (C) the acidic and (D) basic HMWP depletion method, and distribution by percentage for (E) the acidic and (F) basic HMWP depletion method. An overview of the data distribution before and after median normalization and Pearson correlation analysis of biological and technical replicates is provided in the Appendix (FIGURE A-16 and FIGURE A-17). Statistical analysis (two-sided Student's t-test, permutationbased FDR correction, FDR ≤ 0.05, and a Log2 fold change of ±0.485) identified 26 and 30 differentially abundant proteoforms in acidic and basic HMWP depletion experiments, respectively (FIGURE IV-11A and TABLE A-6). In total, 21 proteins were differentially abundant in the acidic HMWP depletion and 24 proteins in the basic HMWP depletion (FIGURE IV-11B). Notably, six proteins had differential abundance changes in both HMWP depletions, with similar abundance changes observed for identical proteoforms or truncated proteoforms of the same protein (TABLE A-7). Quantitative data – The following analysis focused on comparing proteoform-based quantifications from both acidic and basic HMWP depletion methods with the full-proteome peptide-based BUP analysis. While both proteomics-based quantification techniques offer 923 402 175 26 818 350 135 30 211 144 91 21 190 155 78 24 Intern 37% C-term 25%N-term 27% > Intact excised 63% excised > Intact 69% 37% 31% NME: NME: Identified At least one quantification value Quantified Differentially abundant 0 200 400 600 800 1000 ∑ Proteoforms Acidic depletion Basic depletion A Identified At least one quantification value Quantified Differentially abundant 0 50 100 150 200 250 ∑ Protein groups B 2 4 6 8 10 12 14 16 18 20 0 50 100 150 200 ∑ Proteoforms Theo. mass (kDa) C E D F Canonical 12% C-term 31% Intern 22% N-term 35% 11% 2 4 6 8 10 12 14 16 18 20 0 50 100 150 200 C-Term N-Term Internal ∑ Proteoforms Theo. mass (kDa) INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 81 valuable insights, they have different capabilities and limitations. The TDP analysis in this study primarily focused on identifying smaller proteins (<30 kDa), whereas the applied BUP analysis targeted the entire proteome, thus quantifying a broader range of proteins, including larger ones, but it does not identify proteoforms. Detailed variations in quantification, encompassing counts of proteins and proteoforms, are detailed in TABLE A-8 and summarized for both HMWP deletion analysis in TABLE IV-1. TABLE IV-1 | Quantitative Results of Top-down Proteomics and Bottom-up Proteomics. Comparison of proteoformand protein-level quantitative data acquired by full proteome LFQ bottomup and acidic and basic HMWP depletion top-down LFQ analysis. TOP-DOWN QUANTITATIVE RESULTS DIFFERENTIALLY ABUNDANT NOT DIFFERENTIALLY ABUNDANT NOT QUANTIFIED B OTTOM - UP DIFFERENTIALLY ABUNDANT 13 proteins 27 proteins 469 proteins 16 proteoforms 42 proteoforms N/A NOT DIFFERENTIALLY ABUNDANT 22 proteins 67 proteins 1,089 proteins 33 proteoforms 147 proteoforms N/A NOT QUANTIFIED 3 proteins 20 proteins 3 proteoforms 33 proteoforms While 1,558 proteins were quantified in the BUP analysis but not in the TDP analysis, the TDP uniquely quantified a total of 36 proteoforms from 23 proteins that were absent in the BUP analysis (TABLE IV-1). The majority of these (33 proteoforms from 20 proteins) were not differentially abundant in the TDP analysis (TABLE IV-1). However, three proteoforms from three different proteins, an uncharacterized protein (A0A7G5MR36), an acyl carrier protein (A0A7G5MSW4), and a carbohydrate ABC transporter substrate-binding protein (A0A7G5MNS4), showed differential abundance in the TDP analysis (TABLE IV-1). Further differences between the two quantification methods included 27 proteins (represented by 42 proteoforms) which were detected as differentially abundant in BUP dataset but in the TDP dataset (TABLE IV-1). Conversely, 22 proteins (represented by 33 proteoforms) were differentially abundant in the TDP data but not in the BUP data. These included proteoforms of recombination-promoting nuclease RpnA (RpnA; A0A7G5MZ14) and 50S ribosomal protein L29 (RpmC; A0A4P6LZK9), both of higher abundance at pH 7.0, and proteoforms of DNAbinding protein HU (Hup; A0A2S4GGS2) and ribosomal protein S16 (RpsP; A0A4P6M2Y5), which were more abundant at pH 6.0 (FIGURE IV-12). Notably, all differentially abundant proteoforms of DNA-binding protein HU proteoforms exhibited a canonical N-terminal (TABLE A-6), which has been reported to be relevant for forming a DNA–protein complex (Almarza et al., 2015). This complex protects DNA from endonucleolytic cleavage and oxidative stress damage under acidic pH (Almarza et al., 2015). For 67 proteins, both quantification methods detected no significant difference between pH conditions, while 13 proteins were classified as IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 82 differentially abundant by both TDP and BUP (TABLE IV-1). Of these 13 proteins, each TDP depletion method uniquely quantified 5 proteins, with 3 proteins showing differential abundance in both depletion analyses, including the NlpC/P60 domain-containing protein (NlpC; A0A7G5N0P5), D-alanyl-lipoteichoic acid biosynthesis (DltD; A0A7G5MQG2), and a GntR family transcriptional regulator (GntR; A0A7G5MZW5) (FIGURE IV-12 and TABLE IV-2). Notably, both the BUP and the two TDP proteomics-based quantifications revealed identical abundance changes, with GntR being more abundant at pH 7.0 and NlpC and DltD being more abundant at pH 6.0 (FIGURE IV-12 and TABLE IV-2). Notably, even the different proteoforms of GntR exhibited consistent changes in abundance (TABLE IV-2). Moreover, several proteoforms, including the internal 36-amino acid-long proteoform of DltD and the N-terminal truncated 44amino acid-long proteoform of GntR, were quantified as differentially abundant in both TDP datasets (TABLE A-7). FIGURE IV-12 | Comparison Top-down and Bottom-up Label-free-quantification. (A) Bottom-up label-free quantitative analysis. (B) Top-down label-free quantitative analysis from acidic depletion and (C) from basic depletion. Highlighted proteins: D-alanyl-lipoteichoic acid biosynthesis (DltD; A0A7G5MQG2), GntR family transcriptional regulator (GntR; A0A7G5MZW5), DNA-binding protein HU (Hup; A0A2S4GGS2), NlpC/P60 domain-containing protein (NlpC; A0A7G5N0P5), recombinationpromoting nuclease RpnA (RpnA; A0A7G5MZ14), 50S ribosomal protein L29 (RpmC; A0A4P6LZK9) and 30S ribosomal protein S16 (RpsP; A0A4P6M2Y5). GntR RpnA RpmC RpmC Hup Hup NlpC RpsP DltD Hup Hup RpsP 0 1 2 3 4 -9 -6 -3 0 3 6 9 -Log10 p-value Log2(pH 7.0/pH 6.0) ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! RpmC ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! RpnA ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D NlpC D D D D D D D D D D D D D D D D D DltD D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D GntR D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D D 0 1 2 3 4 5 6 7 8 -9 -6 -3 0 3 6 9 -Log10 p-value Log2(pH 7.0/pH 6.0) RpsP Hup q=0.05 A RpmC GntR RpmC RpnA GntR GntR NlpC DltD 0 1 2 3 4 -9 -6 -3 0 3 6 9 -Log10 p-value Log2(pH 7.0/pH 6.0) q=0.05 q=0.05 pH 7.0pH 6.0 B C INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 83 TABLE IV-2 | Overlap of Differentially Abundant Proteins between Bottom-up and Top-down Proteome Analyses. Top-down Log2 ratios represent quantified proteoforms, with multiple values corresponding to different quantified proteoforms. Log2 ratio (pH 7.0/pH 6.0) PROTEIN NAME ACCESSION TDP BUP ACIDIC BASIC NlpC/P60 domain-containing protein A0A7G5N0P5 -6.4 -6.8 -1.3 D-alanyl-lipoteichoic acid biosynthesis A0A7G5MQG2 -2.5 -1.9 -3.2 GntR family transcriptional regulator A0A7G5MZW5 2.3 3.9, 2.3, 1.9 1.6 ABC transporter substrate-binding A0A7G5MSE0 -2.9 - -0.5 Uncharacterized protein A0A7G5N1D2 -2.8 - 1.2 50S ribosomal protein L19 A0A4P6M0Y8 -1.7 - 0.8 XRE family transcriptional regulator A0A7G5N1G9 -1.7 - 1.0 Uncharacterized protein A0A7G5N1C6 2.7 - 3.5 DUF1002 domain-containing protein A0A7G5MNQ5 - -6.0 1.3 Recombinase RecT A0A7G5N1F5 - 0.9 -0.8 NADH peroxidase A0A4P6LUJ8 - 1.0 0.9 Uncharacterized protein A0A7G5N144 - 1.8 2.3 DUF3502 domain-containing protein A0A7G5MQI1 - 7.8 0.8 FIGURE IV-13 | Quantification of the C-terminal Region of the GntR family transcriptional regulator (A0A7G5MZW5). Distribution of identified (A) peptides and (B) proteoforms with highlighted Log2 ratios (pH 7.0/pH 6.0). (C) NCBI-CDD and Alphafold structure prediction revealed an N-terminal HTH (helixturn-helix) DNA-binding domain and C-terminal effector-binding and oligomerization (E-O) domain. 020 40 60 80 100 120 [R].IT...VR.[D] [R].DL...QK.[A] [K].AL...ER.[T] [R].TN...QR.[T] [R].FI...LK.[T] [K].TE...EK.[I] [K].TE...DK.[M] [K].IG...KK.[V] [K].IG...IK.[K] [K].ED...KK.[V] Identified peptides 3.93 2.27 & 2.29 1.85 020 40 60 80 100 120 [M].SW...KK.[-] [L].VY...KK.[-] [L].NM...KK.[-] [N].MI...KK.[-] [N].MI...EE.[K] [N].MI...EE.[E] [N].MI...SE.[E] [N].MI...GL.[S] [N].MI...IG.[L] [M].ID...KK.[-] [I].DN...KK.[-] [D].NL...KK.[-] [N].LK...KK.[-] [L].KT...KK.[-] [T].EL...KK.[-] [L].AS...KK.[-] [G].LS...KK.[-] Sequence length Identified proteoforms HTH Domain NC E-O Domain N-terminal DNA-binding domain C-terminal effector-binding and oligomerization domain C A 1.6 1.0 1.6 1.8 1.4 1.4 B IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 84 3.4 Potential pH-induced Asp-Pro Cleavage The analysis of neo-termini in TDP datasets revealed multiple proteoforms of the ribosomal protein S16 (A0A4P6M2Y5) with Nor C-terminal Asp-Pro peptide bond cleavages. Further sequence analysis of ribosomal protein S16 revealed two Asp-Pro bonds within a loop structure, that separates the prominent N-terminal β-sheets from the C-terminal α-helices (FIGURE IV-14A-B). Overall, five proteoforms exhibiting an Asp-Pro cleavage were identified (FIGURE IV-14C), two of which were differentially higher abundant at pH 6.0 in the acid HMWP depletion TDP analysis (FIGURE IV-12B). Although the basic HMWP depletion TDP analysis also detected these two proteoforms, quantitative data could not be acquired. A re-analysis of the BUP data, incorporating semi-tryptic peptides, identified several peptides with the same Asp-Pro cleavage, most of which were of higher abundance at pH 6.0 (FIGURE IV-14D). Further subcellular localization predictions using Phobius suggested that the resulting proteoforms containing C-terminal α-helices may be directed outside the cytoplasmic space (TABLE IV-3). Additionally, predictions of antimicrobial peptide (AMP) activity using AMPfun suggested that these proteoforms exhibit activity against both Gram-positive and Gramnegative bacteria (TABLE IV-3). FIGURE IV-14 | Asp-Pro Cleavage in 30S Ribosomal Protein S16 (A0A4P6M2Y5). (A) The protein sequence exhibits two Asp-Pro cleavage sites (residues 41-42 and 45-46). (B) The structure prediction highlights the position of the Asp-Pro peptide bonds situated in a structural loop. (C) Identified proteoforms and (D) identified peptides with highlighted Log2 fold changes (pH 7.0/pH 6.0). -0.48 -0.9 -0.7 -1.37 -1.05 -0.50 -0.14 -0.19 -0.22 , -1.04 -0.62 -0.27 -0.10 0.00 , -0.54 , 1.22 -0.38 -1.13 -1.23 -0.54 -1.46 -0.97 -0.31 -0.28 -0.26 -0.76 -1.30 -0.43 -0.84 -0.92 -0.46 -0.91 010 20 30 40 50 60 70 80 [R].MG...YR.[I] [K].KA...YR.[I] [R].II...SR.[S] [R].II...PR.[D] [R].DG...YD.[P] [R].DG...QD.[P] [R].DG...VY.[K] [R].DG...YK.[V] [R].DG...AK.[K] [G].KF...YK.[V] [K].FI...QD.[P] [K].FI...VY.[K] [K].FI...YK.[V] [K].FI...AK.[K] [K].FI...KK.[W] [I].GT...KK.[W] [G].TY...AK.[K] [G].TY...KK.[W] [Y].DP...AK.[K] [D].PN...YK.[V] [D].PN...AK.[K] [D].PN...KK.[W] [Q].DP...KK.[W] [D].PS...AK.[K] [D].PS...KK.[W] [K].VD...SK.[I] [K].KW...SK.[I] [K].WL...SK.[I] [W].LA...SK.[I] [L].AN...SK.[I] [A].NG...SK.[I] [N].GA...SK.[I] [G].AQ...SK.[I] [Q].PT...SK.[I] [K].IF...EK.[-] Sequence length Identified peptides 010 20 30 40 50 60 70 80 [-].MA...NQ.[D] [M].AV...EK.[-] [M].AV...VD.[E] [M].AV...VY.[K] [M].AV...QD.[P] [M].AV...NQ.[D] [M].AV...PN.[Q] [M].AV...YD.[P] [M].AV...TY.[D] [M].AV...GT.[Y] [M].AV...IG.[T] [M].AV...RD.[G] [M].AV...AD.[S] [M].GQ...EK.[-] [G].QK...EK.[-] [G].QK...QD.[P] [Q].KK...EK.[-] [K].KA...EK.[-] [D].SR...EK.[-] [R].DG...EK.[-] [D].GK...EK.[-] [G].KF...EK.[-] [G].TY...EK.[-] [D].PN...EK.[-] [D].PS...EK.[-] [Y].KV...EK.[-] [D].EE...EK.[-] [K].KW...EK.[-] [K].WL...EK.[-] Sequence length Identified proteoforms 30S ribosomal protein S16(A0A4P6M2Y5) 1 MAVKIRLRRM GQKKAPFYRI IVADSRSPRD GKFIEEIGTY DPNQDPSVYK jjVDEEAAKKWL ANGAQPTEVV SKIFKAAGIE K 81 Cleavage sites A C D N C B -1.54 -2.02 Loop INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 85 To investigate the predicted AMP activity, five synthetic peptides, each consisting of 22 amino acids, were designed to cover both the Nand C-terminal Asp-Pro cleavage sites (TABLE A-9). Initial testing of the potential AMP activity was conducted at concentrations ranging from 0.01 to 1 µM against various members of the human gut microbiota (B. thetaiotaomicron, B. longum, A. caccae, C. buytricum, C. ramosum and L. plantrum). These experiments were performed by Kathrin Schäfer (Department of Infectious Diseases and Microbiology, University of Lübeck, UKSH Lübeck; chaired by Prof. Dr. Jan Rupp) and did not reveal significant AMP activity. TABLE IV-3 | Sequence and Structure Predictions of 30S Ribosomal Protein S16. Further investigation was focused on evaluating the possibility of artificially generated Asp-Pro bond hydrolysis during sample preparation, LC-MS measurement, and their potential origin from acidic cultivation. Analysis of the amino acid frequencies surrounding non-tryptic cleavage sites in the BUP datasets revealed distinct sequence logos among different bacterial species (FIGURE IV-15). Specifically, B. thetaiotaomicron and B. producta showed an increased relative frequency of Asp at the P1 position and Pro at the P1´ position (FIGURE IV-15A). Interestingly, this frequency increased for peptides identified at pH 6.0, particularly for B. thetaiotaomicron, suggesting that acidic conditions may promote the potential hydrolysis of Asp-Pro bonds. In contrast, B. longum exhibited a significantly increased relative frequency only for Pro at the P1´ position. Further analysis focusing on peptides with Asp at the P1 position (FIGURE IV-15C) or Pro at the P1´ position (FIGURE IV-15D) confirmed the presence of peptides with Asp-Pro cleavage and also indicated their occurrence for B. longum. S EQUENCE FRAGMENT (#AA) AMP (P REDICTION ) AMP T ARGET (P REDICTION) L OCALIZATION (P REDICTION) S TRUCTURE P REDICTION First DP motif cleaved [M].AV...YD.[P] N -term ( 40) AMP (0.757) Gram -positive (0.550), Gram -negative (0.578) Cytoplasmic ( 0.775) [D].PN...EK.[ -] C -term (40) AMP (0.702) Gram -negative (0.586) Non -cytoplasmic ( 0.681) Second DP motif cleaved [M].AV...QD.[P] N -Term ( 44) AMP (0.7181) Gram -negative (0.527) Cytoplasmic ( 0.774) [D].PS...EK.[ -] C -Term (36) AMP (0.788) Gram -positive (0.575), Gram -negative (0.566) Non -cytoplasmic ( 0.685) IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 86 FIGURE IV-15 | Bottom-up Peptide Cleavage Analysis Highlights Asp-Pro Sequence Logos. IceLogo plot illustrating the relative frequency of P2-P2′ amino acids of identified peptides for B. thetaiotaomicron, B. producta and B. longum. (A) Identified peptides at pH 7.0 and (B) at pH 6.0. (C) Identified peptides with Asp at the P1 position and (D) with Pro at the P1´ position. Amino acids that are significantly enriched (top) or depleted (bottom) (p ≤ 0.05) are colored black, with Asp and Pro residues highlighted in red. Moreover, significantly higher abundance at pH 6.0 was observed for peptides with cleaved Nor C-terminal Asp-Pro and Asp-Xaa bonds in both B. thetaiotaomicron (FIGURE IV-16A) and B. producta (FIGURE IV-16B). Conversely, peptides with intact Asp-Pro showed no significant changes in abundance and were distributed around zero (FIGURE IV-16A-B). Additionally, at pH 7.0, peptides with cleaved Nor C-terminal Asp-Xaa bonds exhibited significantly higher abundance in B. longum (FIGURE IV-16C). FIGURE IV-16 | Peptide Abundance Distribution. Distribution of all identified peptides, semitryptic peptides (excluding those with intact Asp-Pro, cleaved Asp-Pro and Asp-Xaa bonds), peptides with intact Asp-Pro bonds, peptides with cleaved Asp-Pro and cleaved Asp-Xaa bonds. (A) B. thetaiotaomicron, (B) B producta, and (C) B longum. Significant differences were calculated via oneway ANOVA with Dunnett’s correction. * (p < 0.05); ** (p < 0.01); **** (p < 0.0001). The number on the top indicates the number of identified peptides. P2 P1 P1' P2' % Difference -45 -22.5 22.5 45 ∑ Analyzed sites = 1910 P2 P1 P1'P2' % Difference -30 -15 15 30 ∑ Analyzed sites = 2778 A P2 P1 P1'P2' % Difference -45 -22.5 22.5 45 ∑ Analyzed sites = 2088 pH 7.0 pH 7.0 pH 7.0 B. longum B. thetaiotaomicron B. producta P2 P1 P1' P2' % Difference -100 -50 50 100 ∑ Analyzed sites = 377 C P2 P1 P1' P2' % Difference -100 -50 50 100 ∑ Analyzed sites = 220 P2 P1 P1' P2' % Difference -100 -50 50 100 ∑ Analyzed sites = 131 Asp P1 Asp P1 Asp P1 B. longum B. thetaiotaomicron B. producta P2 P1 P1' P2' % Difference -45 -22.5 22.5 45 ∑ Analyzed sites = 1924 P2 P1 P1'P2' % Difference -30 -15 15 30 ∑ Analyzed sites = 2795 B P2 P1 P1'P2' % Difference -45 -22.5 22.5 45 ∑ Analyzed sites = 2500 pH 6.0 pH 6.0 pH 6.0 B. longum B. thetaiotaomicron B. producta P2 P1 P1' P2' % Difference -100 -50 50 100 ∑ Analyzed sites = 223 D P2 P1 P1' P2' % Difference -100 -50 50 100 ∑ Analyzed sites = 172 P2 P1 P1' P2' % Difference -100 -50 50 100 ∑ Analyzed sites = 169 Pro P1' Pro P1' Pro P1' B. longum B. thetaiotaomicron B. producta 25490 1976 752 127 257 -3 -2 -1 0 1 2 3 4 Log2 (pH 7.0/pH 6.0) * **** A ** 18164 2126 574 101 99 -3 -2 -1 0 1 2 3 4 Log2 (pH 7.0/pH 6.0) **** B **** **** ** **** 9247 1753 457 58 29 -3 -2 -1 0 1 2 3 4 Asp-Xaa bonds cleaved Asp-Pro bonds cleaved All identified peptides Asp-Pro bonds intact Semitryptic peptides Log2 (pH 7.0/pH 6.0) C INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 87 3.5 Phosphorylation of HPr proteins An open-modification search of the TDP dataset identified several potential serylphosphorylations on histidine-containing phosphocarrier proteins (HPr). Overall, a high number of matching fragment ions covering the phosphorylation site, along with a matching distribution of theoretical and experimentally observed peaks of unmodified and phosphorylated HPr proteoforms, was observed (FIGURE IV-17 and FIGURE A-19A-D). Using ProSight Annotator, potential phosphorylation sites were incorporated into HPr protein entries, facilitating a variable search for HPr phosphorylation. Combined with an additional database search of the bottom-up proteomic data that included phosphorylation as a variable modification at serine, histidine, threonine, tyrosine, and arginine residues, several phosphorylated peptides and proteoforms were identified. These findings suggested Ser-41 phosphorylation for HPr (A0A2S4GRU0 and A0A7G5MYS7), Ser-46 phosphorylation for HPr (A0A4V0Z7D2), and Arg-46 phosphorylation for HPr (A0A7G5MP53) (FIGURE IV-18). Although peptides spanning the respective phosphorylation sites of HPr (A0A7G5MP53 and A0A4V0Z7D2) were detected by the bottom-up proteomics analysis, phosphopeptides could only be identified for two HPr variants (A0A2S4GRU0 and A0A7G5MYS7) (FIGURE A-19E-H). FIGURE IV-17 | Phosphoserine Modification of HPr (A0A2S4GRU0). Distribution of experimentally observed and theoretical peaks of identified proteoforms for (A) unmodified HPr and (B) phosphorylated HPr. (C) Identified proteoform sequence and fragment ion spectrum, highlighting the detected phosphorylation site at Ser-41 in yellow. The identified band y-ions are highlighted in blue and red, respectively. Dot plot illustrates the mass error in ppm for each observed band y-ion, respectively. Error (ppm) Relative Abundance (%) m/z A B C IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 88 FIGURE IV-18 | Sequence Alignment of HPr Proteins Highlighting Proposed Phosphorylation Sites. Positions of the two conserved amino acid residues: a histidine (His-15) near the N-terminus and a serine (Ser-46) in the central part of the protein, are highlighted. The quantitative data of the BUP analysis revealed a differential abundance of HPr (A0A2S4GRU0) at pH 6.0. Moreover, analysis of peptide-level abundance revealed a higher abundance of phosphopeptides at pH 6.0, while unmodified peptides were more abundant at pH 7.0 (FIGURE IV-19A-B). Statistical analysis of the TDP data, which incorporated the variable HPr phosphorylation sites, quantified several phosphorylated proteoforms as differentially abundant at pH 6.0, both by the acidic (FIGURE IV-19C) and the basic HMWP depletion (FIGURE IV-19D). Conversely, unmodified or N-terminal formylated proteoforms did not exhibit differential abundance (FIGURE IV-19C-D). Overall, the observed increase in the abundance of phosphorylated peptides and proteoforms of HPr (A0A2S4GRU0) at pH 6.0 suggests that HPr phosphorylation may represent a direct response to culture acidification. FIGURE IV-19 | Quantification of Phosphorylation on HPr (A0A2S4GRU0). (A) Scaled peptide abundance at pH 6.0 and pH 7.0. Peptides with serine phosphorylation (P) or methionine oxidation (O) are highlighted. One boxplot is absent as the phosphopeptide was exclusively identified at pH 6.0. Proteoform-directed label-free analysis from (B) acidic and (C) basic depletion. Phosphorylated HPr proteoforms with differential abundance at pH 6.0 are highlighted in yellow (Two-sided Student's t-test, Permutation-based FDR corrected, q-value of 0.05). Identified using A0A4P6M732 8 -VNNLIGLHLRPAG-20 42-ANAKSVLSV-50 A0A7G5MP53 8 -ITDPEGIHARPAG—20 42-GDCKRIFGI-50 Top -down proteomics A0A4V0Z7D2 8 -IGISNGLEARPIA—20 42-VNAKSIMGM-50 Top -down proteomics A0A7G5MYS7 8 -LNET-----GDVK—20 37-IDAKSILGV-45 Top -down & Bottomup proteomics A0A2S4GRU0 8 -LNSI-----DKVK—20 37-IDAKSIMGI-45 Top -down & Bottomup proteomics A0A4P6LXB3 8 -FSEI-----NEIK—20 37-IDAKSILGM-45 A0A4P6LTJ7 7 -FKEV-----DEIV—19 36-VDAKSIMGM-44 P P ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! Unmodified N-Formyl HPr(Ser-P) ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! !! HPr(Ser-P) 0 1 2 3 4 -9 -6 -3 0 3 6 9 -Log10 p-value q=0.05 ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! Unmodified N-Formyl HPr(Ser-P) ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! HPr(Ser-P) HPr(Ser-P) 0 1 2 3 4 -9 -6 -3 0 3 6 9 -Log10 p-value q=0.05 [K].IS...VK.[S] [K].IS...DK.[V] [K].SF...GR.[Y] [K].SF...AK.[S] [K].SF...TK.[F] [K].FD...GR.[Y] [R].YV...SK.[A] [K].SI...SK.[A] [K].SI...SK.[A] [K].SI...SK.[A] [K].SI...SK.[A] [K].AI...VD.[-] 0 1 2 3 1 >50 PSMs 40 30 20 10 Scaled abundance pH 6.0 pH 7.0 A B0 20 40 60 80 O P Sequence position P P OO D C Log2(pH 7.0/pH 6.0) Log2(pH 7.0/pH 6.0) INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 95 synthesis of histidine, which consumes other amino acids and metabolites, can be disadvantageous during acid stress conditions. Therefore, it can be hypothesized that B. producta also possesses a histidine uptake system to assimilate extracellular histidine, to avoid the high energy cost associated with amino acid biosynthesis. Although histidine uptake in B. producta has not been studied to my knowledge, studies in other Gram-positive bacteria suggest the adaptation of common ABC transporters for histidine uptake (Vitreschak et al., 2008). To experimentally validate the importance of histidine biosynthesis or the uptake of histidine in the acid stress response of B. producta, several follow-up experiments can be performed. For instance, transposon sequencing can be used to compare the bacterial fitness of wild-type B. producta against a single-gene disruption library grown in media with and without histidine supplementation under acid stress. This approach can identify genes required or detrimental for the growth of B. producta under acid stress conditions. While single knockout studies in E. coli suggested that most amino acid transport and metabolism genes, including histidine biosynthetic genes, are non-essential (Baba et al., 2006), transposon sequencing studies in S. aureus revealed a crucial gene (SAUSA300_0846) encoding a histidine transporter (Beetham et al., 2024). The histidine transporter function was confirmed by measuring the uptake of radio-labeled histidine in both wild-type S. aureus and the mutant strain, where histidine uptake was largely lost in the mutant. In the absence of exogenous histidine, the wild-type strain exhibited an additional 5-hour lag phase, indicating its reliance on histidine for growth and its ability to adapt and synthesize histidine. Growth experiments conducted at pH 4.3 and pH 7.2, with or without histidine, confirmed the importance of histidine transport via the SAUSA300_0846 transporter for S. aureus growth at acidic pH. Notably, the mutant strain exhibited a more than 200-fold increase in the expression of histidine biosynthesis genes under acid stress conditions. Currently, to the best of my knowledge, their study stands as the pioneering exploration of a potential histidine-dependent acid tolerance system. Their meticulous validation of the potential involvement of histidine in the acid resistance of S. aureus sets an excellent example and should serve as a benchmark for future research into the potential histidine-dependent acid tolerance system of B. producta. Top-down Proteomic Analysis – In addition to the conventional proteome analysis using a BUP approach, a TDP approach was employed to characterize and quantify individual proteoforms. To enhance the detection of the low-molecular-weight proteome, two depletion methods and LC-FAIMS-MS2 analysis with internal CV stepping (Kaulich et al., 2022a) were utilized to boost sensitivity and increase the number of proteoforms. Recent studies have shown that utilizing FAIMS with internal CV stepping not only doubles the number of quantified proteoforms but also maintains quantification accuracy (Kline et al., 2023). While a newer version of Proteome Discoverer (v.3.0) was used by Kline and colleagues, allowing for the IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 96 specification of multiple CVs per file in the spectrum selector node, database searches in this study were conducted using Proteome Discoverer (v.2.5.0.400). This version does not support multiple CVs per file. Consequently, the resulting data files had to be sliced into subsets filtered by the applied CVs before database searching (Leipert et al., 2023). The TDP analysis of acidic and basic HMWP depletions of B. producta quantified 175 proteoforms (19% of the 923 identified proteoforms) and 135 proteoforms (17% of the 818 identified proteoforms), respectively (FIGURE IV-11A and B). The low number of quantified proteoforms can be attributed to the significant number of missing values observed (38% and 42%), which is comparable to other TDP-LFQ studies that reported 43% (Leipert et al., 2023) or 54.6% (Ntai et al., 2016). Interestingly, the number of missing values can depend on the applied CVs, which ranged in the proteoform analysis of Caenorhabditis elegans from 23% (CV -20) to 43% (CV -50) (Leipert et al., 2023). Notably, Kline and colleagues quantified 499 E. coli proteoforms, which represents only 29% of the 1,719 identified proteoforms, potentially indicating a similar occurrence of missing values in their analysis (Kline et al., 2023). The higher number of missing values in a proteoform-centric analysis can be attributed to broader isotopic envelopes and wider charge state ranges compared to a peptide-centric analysis (Basharat et al., 2023). Consequently, co-eluting proteoforms of different charge states can share overlapping m/z ranges even when they have highly distinct masses. This phenomenon can interfere with resolving and accurately quantifying proteoforms by spectral deconvolution algorithms, such as Xtract deconvolution employed by the ProSightPD node in Proteome Discoverer. While multi-dimensional protein fractionation strategies can mitigate proteoform co-elution (Cassidy et al., 2021a; Kaulich et al., 2024), the increased complexity can pose challenges for LFQ analysis in correctly determining proteoform abundances across numerous fractions. Therefore, future advancements in deconvolution algorithms are essential, particularly those that possess the capability to resolve and quantify co-eluting proteoforms or acquire precursors with minimal interference. Such advancements are crucial for mitigating data loss and ensuring the comprehensive characterization of proteomic samples, ultimately advancing TDP-LFQ analysis (Jeong et al., 2020, 2022; Basharat et al., 2023). Although the TDP analysis primarily targeted the LMWP (<30 kDa), it detected differential abundance of 38 proteins (TABLE IV-1). Specifically, 21 proteins exhibited differential abundance in the acidic depletion analysis and 24 in the basic depletion analysis (FIGURE IV11B). Six proteins showed differential abundance in both depletion methods (TABLE A-7). Importantly, both depletion methods allowed comparable quantification of abundance changes for identical proteoforms of the same protein (TABLE A-7). A comparison of the quantification results of BUP and TDP revealed the non-differential abundance of 67 proteins and the differential abundance of 13 proteins (TABLE IV-1). Among these 13 proteins, three (GntR; A0A7G5MZW5, NlpC; A0A7G5N0P5, and DltD; A0A7G5MQG2) exhibited differential abundance in both TDP depletion analyses and the BUP INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 97 analysis (FIGURE IV-12 and TABLE IV-2). Notably, at pH 6.0, the D-alanyl-lipoteichoic acid biosynthesis protein (DltD; A0A7G5MQG2), responsible for catalyzing the D-alanylation of lipoteichoic acid, exhibited increased abundance. D-alanylation of teichoic acids can mask the negative charge of the cell membrane, enhancing growth and survival in low pH environments (Boyd et al., 2000; Wu et al., 2022). While BUP quantification aims to measure the abundance of peptides derived from a given protein, it often fails to capture proteoform variations due to limitations in peptide-to-protein inference. Estimations by Ntai and colleagues suggest that around 40% of proteoform-level dynamics in abundant, low-molecular-weight (<30 kDa) proteins remain undetected by BUP (Ntai et al., 2016). Despite these limitations, both peptideand proteoform-level quantification of GntR indicate a higher abundance at pH 7.0 in the C-terminal protein region (FIGURE IV-13). Similarly, BUP analysis of the DNA-binding protein HU (Hup; A0A2S4GGS2) indicated higher abundance of several peptides from the N-terminal region at pH 6.0, potentially crucial for DNA protection under acidic stress (Almarza et al., 2015). However, not all peptides were quantified, and some also showed increased abundance at pH 7.0. In contrast, proteoform analysis revealed a differentially increase in abundance of DNA-binding protein HU proteoforms spanning the N-terminal region at pH 6.0 (TABLE A-6). This further emphasizes the importance of integrating TDP analyses to identify potential biologically proteoform abundances resulting from variant expression, proteolytic truncation, or changes in PTM stoichiometry. Interestingly, many of the differentially abundant proteoforms from the same protein had abundance changes primarily characterized by truncation, rather than other PTMs (TABLE A6). However, the protein sequence database of B. producta used for this analysis contained only a few automatically UniProt-annotated PTMs, with most proteins lacking information on potential or validated modifications. Consequently, during the Proteome Discoverer database search, proteoforms with potential PTMs that were not annotated in the UniProt database remained elusive. Therefore, an additional discovery-open modification search was utilized to identify potential PTM-carrying proteoforms. This approach successfully identified and quantified multiple seryl-phosphorylations on HPr proteins (FIGURE IV-18 and FIGURE IV-19). Although these findings require further validation, the application of a discovery-open modification search may be considered to complement future analysis, especially for bacterial species with limited UniProt-annotated PTMs. Overall, the combination of BUP and TDP analysis identified potential seryl-phosphorylations on HPr (A0A2S4GRU0, A0A7G5MYS7, and A0A4V0Z7D2), along with one arginylphosphorylation on HPr (A0A7G5MP53) (FIGURE IV-18). The presence of phosphopeptides with missed cleavages near the proteolytic site (P1´) (FIGURE A-19E and G) could potentially strengthen the identification of phosphorylation sites, as nearby phosphorylation sites (P1´, P2', and P3') can interfere with tryptic cleavage (Dickhut et al., 2014; Gershon, 2014). Although the arginyl-phosphorylation was detected with 68 PrSMs, the absence of fragment ions around IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 98 the modification site, the lack of detected phosphopeptides, and the limited literature on arginyl-phosphorylation in HPr proteins make its PTM localization uncertain with the currently available data. Moreover, sequence analysis of B. producta's HPr proteins revealed notable differences from other members of this protein family. While typical HPr proteins feature two highly conserved amino acid residues, a histidine (His-15) near the N-terminus and a serine (Ser-46) in the central part of the protein, serving as phosphoryl group acceptors (Meadow et al., 1990; Brochu and Vadeboncoeur, 1999; Casabon et al., 2006), the majority of B. producta's HPr proteins lack five amino acids near the N-terminus, including the conserved His-15. This deficiency potentially makes them incapable of forming doubly phosphorylated HPr(His∼P)(Ser∼P) (FIGURE IV-18). Typically, HPr phosphocarrier proteins are involved in carbohydrate phosphorylation during transport into bacterial cells via the phosphotransferase system (Meadow et al., 1990). Depending on their phosphorylation state, HPr proteins can also act as coregulators of the catabolite global regulator CcpA (Homeyer et al., 2007), which controls the catabolite repression/activation of up to 10% of total genes (Poncet et al., 2004). This can provide a direct link to the metabolic state of the bacterial cell and a regulatory role in the quorum sensing of the cell (Ha et al., 2018). The combined quantitative data of the BUP and the TDP analysis showed that the HPr protein (A0A2S4GRU0), their phosphorylated peptides as well as their phosphorylated proteoforms were differentially more abundant at pH 6.0 (FIGURE IV-19). In contrast, unmodified peptides covering the same protein sequence and unmodified or formylated proteoforms were higher abundant at pH 7.0 (FIGURE IV-19). Notably, this observation was independent of the applied depletion method for the TDP analysis (FIGURE IV19C-D). The observed increase in HPr abundance and seryl-phosphorylation under acidic growth conditions aligns with prior research demonstrating a similar trend of increased HPr protein abundance and HPr(Ser-P) formation in response to culture acidification (Casabon et al., 2006; Heunis et al., 2014). Hence, the results of this study further emphasize the significant influence of culture pH on the phosphorylation status of HPr proteins. 4.4 Asp-Pro peptide bond hydrolysis Notably, peptides with Nor C-terminal Asp-Pro bond cleavages were detected for all three bacteria (FIGURE IV-15C-D). Specifically, for B. thetaiotaomicron, there was an increased relative frequency of Asp at the P1 position and Pro at the P1´ position for peptides identified at pH 6.0 (FIGURE IV-15B), suggesting that acidic conditions may promote the hydrolysis of Asp-Pro bonds. This potential biological response is further supported by a significant increase in peptides exhibiting Nor C-terminal Asp-Pro and Asp-Xaa cleavages for B. thetaiotaomicron (FIGURE IV-16A) and B. producta (FIGURE IV-16B). In contrast, the proteome of B. longum did INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 99 not show an Asp-Pro sequence logo (FIGURE IV-15A-B) and neither significant changes in the abundance of peptides with Nor C-terminal Asp-Pro cleavages (FIGURE IV-16C). Several factors, such as the peptide sequence, protein folding, temperature, and pH value, can affect the hydrolysis of peptide bonds (Marcus, 1985; Li et al., 2009). The Asp-Pro bond, in particular, exhibits increased susceptibility to peptide bond hydrolysis under acidic conditions and elevated temperatures (Marcus, 1985; Li et al., 2009). The proposed acidolysis of Asp-Pro bonds suggests that the β-carboxyl group of Asp initiates a nucleophilic attack on the carbonyl carbon of the amide bond, facilitating intraresidue cyclization, forming an unstable cyclic anhydride intermediate, which may cause a break in the polypeptide chain (Piszkiewicz et al., 1970). This reaction requires the presence of an adjacent protonated amide nitrogen. Generally, the enhanced rate of Asp-Pro cleavage can be attributed to the greater basicity of the nitrogen atom as part of proline's cyclic structure, which increases its basicity (pKa of 10.6) compared to the primary amine groups of other amino acids (pKa of 8.7 to 9.9), thus increasing its nucleophilicity and facilitating faster hydrolysis of Asp-Pro bonds under acidic conditions (Piszkiewicz et al., 1970). The importance of such a labile peptide bond can be crucial for planning and performing proteomic sample preparation and subsequent LC-MS measurement, especially for Nand Cterminomics experiments. For example, acidification using pure TFA for cell lysis and high concentrations of Tris for neutralization such as applied in the SPEED protocol, should be avoided (Doellinger et al., 2020). Furthermore, Tris-based buffers exhibit significant pH changes upon temperature change compared to other buffer systems such as sodium phosphate buffer solutions, which are generally more resistant to acidification upon temperature change (Kolhe et al., 2010). The choice of reducing agent can also be critical, as TCEP can cause a significant drop in pH compared to dithiothreitol (DTT) (Scheerlinck et al., 2015). Additionally, different buffers are required depending on the protease used for digestion. For example, pepsin or neoprosin necessitates highly acidic conditions (pH 1.5 to pH 2.5) for optimal enzymatic activity, a condition that has been demonstrated to significantly increase the hydrolysis of Asp-Pro bonds (Schräder et al., 2017). Moreover, most in-solution workflows involve SPE for sample clean-up, typically involving organic solvents with TFA or FA, which can also increase Asp-Pro bond cleavage (Winkels et al., 2022). Generally, prolonged exposure to acidic conditions or heating should be avoided (Kaulich et al., 2024), and gentler methods such as freeze-drying or desiccation to evaporate organic solvents after SPE are preferable. Alternatively, different sample and clean-up protocols, such as SP3 (Hughes et al., 2019) or FASP (Manza et al., 2005; Wiśniewski et al., 2009), can be applied, which often do not require an additional SPE step. Typically, proteomic LC-MS analysis usually employs an acidic mobile phase and a thermostat-controlled column oven with elevated temperatures for optimal chromatographic separation, controlled retention times, and reduced pressure of the LC system (García, 2005; Lenčo et al., 2022). However, these conditions, in combination with IV | INFLUENCE OF PH ON BACTERIAL PROTEOMES P A G E | 100 prolonged in-column residence time, can induce in-column peptide hydrolysis of Asp-Pro bonds (Lenčo et al., 2021). Despite the significant enrichment and abundance of peptides with Asp-Pro and Asp-Xaa cleavage at pH 6.0 for B. thetaiotaomicron and B. producta, the possibility of artificial peptide bond hydrolysis during sample preparation and LC-MS analysis cannot be ruled out. While employing the same sample preparation steps and LC-MS setup for the full proteome analysis of the three HGM bacteria, slight variations, such as the total sample volume after SPE and extended vacuum centrifuge evaporation times after SPE sample clean-up, may have occurred. Additionally, the potential for the observed results to be generated in vivo, reflecting a biological effect of acidic cultivation, cannot be disregarded. Further experiments using a proteomic sample preparation workflow designed to limit the artificial generation of Asp-Pro cleavage in combination with an N-terminomics approach may clarify the origin of these peptide formations (Winkels et al., 2022). While Asp-Pro cleavage is reported to be artificial in most studies, selective cleavage of AspPro bonds has been reported to be required for correct folding (Patrick and Egland, 2019), intermolecular protein cross-linking forming an Asp-Lys isopeptide bond (Osička et al., 2004) or covalently linking chondroitin sulfate forming a protein-glycosaminoglycan-protein complex (Zhuo et al., 2004). Moreover, several studies have reported in vivo cell-compartment and pHspecific autocatalytic cleavage between Asp-Pro bonds of several proteins exhibiting a GlyAsp-Pro-His (GDPH) sequence, which is commonly found in von Willebrand factor type D domains. For example, autocatalytic cleavage at body temperature has been reported for several human proteins before being secreted into the intestinal mucus. This includes GDPH cleavage of IgGFc-binding protein (FCGBP) or MUC2 mucin, which is triggered by the lower pH of the endoplasmic reticulum or the Golgi apparatus, respectively (Lidell et al., 2003; Ehrencrona et al., 2021). Similarly, autocatalytic GDPH cleavage of proH3 precursors occurs during passage through the low pH of the Golgi complex (Thuveson and Fries, 2000). While the exact function of GDPH cleavage is still a topic of ongoing investigations, it may be an important factor for covalent protein or carbohydrate cross-linking, especially in the mucus (Lidell et al., 2003). For instance, alterations in the interactome of proteins within the arterial extracellular matrix, potentially contributing to protein aggregation cascades, are suggested by results from the enrichment of Asp-Pro cleavage products of NOTCH3, associated with disease-affected brain tissue (Lee et al., 2023). The results of this study indicate a potential acidic-induced Asp-Pro cleavage of 30S ribosomal protein S16 (A0A4P6M2Y5; FIGURE IV-14), with Cand N-terminal fragments exhibiting potential AMP activities (TABLE A-9). Although preliminary AMP testing against members of the HGM did not confirm such activity, further experiments with other members of the gut microbiota may be necessary to determine the presence of potential AMP activity. INFLUENCE OF PH ON BACTERIAL PROTEOMES | IV P A G E | 101 Similar to the proteoform concept, which suggests that different protein species obtained through processes like truncation or other post-translational modifications can have distinct functions (Jungblut et al., 2016), several ribosomal proteins have been identified with "moonlighting" functions, wherein a single protein resulting from one gene serves multiple roles in a cell or organism, including exhibiting antimicrobial activities (Hurtado-Rios et al., 2022). For example, Bacillus tequilensis isolated from healthy human feces can secrete ribosomal protein L1 as an antimicrobial molecule (Ghoreishi et al., 2023), while Lactobacillus salivarius, isolated from the feces of four-month-old human infants, secretes ribosomal proteins L27 and L30 with antimicrobial activities (Pidutti et al., 2018). Notably, several antimicrobial peptides from ribosomal proteins L30, L39, S19, and S30 have been identified in cellular extracts of human non-inflamed colonic mucosa, indicating a diverse presence of ribosomal antimicrobial peptides in human colonic mucus (Howell et al., 2003; Tollin et al., 2003; Antoni et al., 2013). Whether B. producta ribosomal protein S16 may function as a moonlighting protein with antimicrobial activity remains to be elucidated. P A G E | 102 P A G E | 103 V PROTEOGENOMIC ANALYSIS OF B. PRODUCTA 1 Introduction and Summary ........................................................................................... 105 2 Experimental Design ..................................................................................................... 107 2.1 Culture Condition Effects on SEP Production ....................................................... 107 2.2 Peptide and Proteoform Validation ........................................................................ 108 3 Results ............................................................................................................................ 110 3.1 Identification and Validation of Translated SEP .................................................... 110 3.2 SEP Proteoform Diversity ...................................................................................... 113 3.3 Impact of Cultivation Conditions on SEP Translation ............................................ 115 3.4 Biochemical Predictions of Identified SEP ............................................................ 116 4 Discussion and Conclusion .......................................................................................... 117 4.1 Challenges and Advances in SEP Identification ................................................... 117 4.2 Future Directions of SEP Research ...................................................................... 119 V | PROTEOGENOMIC ANALYSIS OF B. PRODUCTA P A G E | 104 Parts of the following chapter have been published in “Identification of proteoforms of short open reading frame-encoded peptides in Blautia producta under different cultivation conditions.” Genth et al., Microbiology Spectrum, 11(6), e0252823, (2023). Supplementary material for (Genth et al., 2023) Additional supplementary information’s is freely available for download at the publisher’s website https://doi.org/10.1128/spectrum.02528-23. The MS proteomics raw data and complete Proteome Discover search results have been deposited to the ProteomeXchange Consortium (http://www.proteomexchange.org/) via the PRIDE (Vizcaíno et al., 2014) partner repository with the data set identifier PXD041979. PROTEOGENOMIC ANALYSIS OF B. PRODUCTA | V P A G E | 111 Peptide Validation – The applied BUP approach identified 2,379 proteins across all seven cultivation conditions, of which 261 proteins (9.1%) had less than 100 amino acids (FIGURE V2). Among these, 55 proteins passed the PSM filtering criteria, represented by 142 peptides, which were subsequently subjected to the verification process (FIGURE V-2). The validation of these non-canonical peptides using PepQuery led to 107 peptides passing the strict verification criteria (FIGURE V-3). Out of the 35 peptides that did not pass the PepQuery verification process, 12 were not matched by PepQuery to any MS/MS spectra with sufficient quality scores, 4 had superior matches to peptides in the reference database, and 19 were either better matched with reference peptides carrying potential PTMs or did not meet statistical thresholds (p-value < 0.01) matching the non-canonical sequence (FIGURE V-3). Subsequent BLASTp analysis of the remaining peptides that successfully passed PepQuery identified a single amino acid variation (SAV) in one peptide compared to a RefSeq protein. Considering the potential occurrence of single nucleotide polymorphisms in genetic sequences, which can arise through various mechanisms such as DNA replication errors, there was an increased likelihood that the detected sequence variant may not necessarily be linked to a novel SEP. To ensure the accuracy and reliability of the results, both the peptide and its corresponding SEP were consequently excluded as a precautionary measure. Comparison with different in silico digestions, even when treating leucine and isoleucine as equivalent in peptide sequence analysis (FIGURE A-20), led to the exclusion of two additional peptides, but not their respective SEP. Therefore, potential peptide contamination due to the use of protein-containing yeast extract or BHI could be largely excluded. In conclusion, 104 peptides were validated (FIGURE V-3), which led to the identification of 44 SEP (FIGURE V-2). FIGURE V-3 | Validation of Non-canonical Peptides. Overview of stringent filtering and PepQuery peptide validation. Abbreviation: SAV (single amino acid variation). V | PROTEOGENOMIC ANALYSIS OF B. PRODUCTA P A G E | 112 Proteoform Validation – A total of 1,572 proteoforms were identified and subjected to a rigorous filtering and validation process, resulting in the identification of 177 potential noncanonical proteoforms (FIGURE V-4). Among these, 60 proteoforms met the criteria for both protein size and number of PrSMs. Additional filtering, which included specific criteria such as C-Score, E-value, and the number of fragment ions, led to the exclusion of 5 proteoforms that did not meet the established criteria (FIGURE V-4). As a result, a total of 55 validated proteoforms (FIGURE V-4), derived from 19 SEP, successfully passed the filtering and validation process (FIGURE V-2). FIGURE V-4 | Validation of Non-canonical Proteoforms. Stringent filtering and proteoform validation, including the distribution of identified Nand C-termini and post-translational modifications (PTMs) on SEP proteoforms. The average number of PrSMs identified for each SEP proteoform was 24, with some exhibiting considerably higher numbers, up to 2,360 (FIGURE V-5A). The identified SEP proteoforms had a median -Log E-value of 17 (FIGURE V-5B), 26 fragment ions (FIGURE V-5C), and a residue cleavage rate of 44% (FIGURE V-5D). These results, combined with a median sequence coverage of 98% (FIGURE V-5E), collectively provide strong evidence for a high confidence level associated with the identified proteoforms. FIGURE V-5| Metrics of Proteoform Identifications. (A) Number of PrSMs, (B) -Log E-values, (C) matching fragment ions, (D) residue cleavage, and (E) sequence coverage of proteoform identifications. The box-and-whisker plots illustrate the lower quartile and upper quartile, with the median displayed as a horizontal line and the mean depicted as a cross. Whiskers represent the minimum and maximum values that fall within 1.5 times the interquartile range 0 50 100 150 200 250 1500 3000 Number of PrSMs A 0 20 40 60 80 100 Matching fragment ions C 0 20 40 60 80 100 % Residue cleavage D 0 20 40 60 80 100 % Sequence coverage E 0 5 10 15 20 25 -Log E-value B PROTEOGENOMIC ANALYSIS OF B. PRODUCTA | V P A G E | 113 3.2 SEP Proteoform Diversity Neo-termini analysis of the identified proteoforms identified 26 canonical proteoforms of which 16 retained their N-terminal methionine and 10 lacked their N-terminal methionine, indicating N-terminal methionine excision (NME) (FIGURE V-4). The remaining proteoforms exhibited either Nor C-terminal truncations, or truncations at both ends (FIGURE V-4). Analysis of the positions surrounding neo-termini sites (P2-P2') indicated an increased specificity for methionine at the P1 position (FIGURE V-6A). While proteoforms resulting from NME were excluded from this analysis, these results may still indicate potential NME by methionine aminopeptidase. The identification of N-terminally methionine-truncated proteoforms, in conjunction with the presence of smaller amino acid residues, such as serine or alanine in the P1´ position – known for their increased efficiency in methionine removal (Meinnel et al., 1993) – suggests the possibility of alternative translation initiation. Such proteoform variants were observed for BP4 and BP7, with N-terminal truncation initiating at methionine position 5, as well as canonical variants with or without N-formylmethionine (FIGURE V-7). Alternative initiation seems likely, given the short N-terminal portion, which is insufficient for a signal peptide (Peng et al., 2019). Similarly, BP14 proteoforms lacked the encoded N-terminal portion of 24 amino acids, including a second methionine at position 24 (FIGURE V-6B). The missing N-terminal portion fell within the typical range of a potential signal peptide (Peng et al., 2019). Phobius analysis (Käll et al., 2004) revealed a potential signal peptide for the missing N-terminal segment (posterior probability: 0.87) and a non-cytoplasmic region for the remaining C-terminal segment (posterior probability: 0.96) (FIGURE V-6B). Whether this represents alternative initiation, signal peptidase cleavage, or another truncation event remains to be elucidated. Understanding these mechanisms and exploring whether these variants play specialized roles in response to specific environmental conditions or stimuli could provide insights into the regulatory processes controlling their expression. FIGURE V-6 | Neo-termini Analysis of Identified SEP Proteoforms. (A) IceLogo analysis of neo-termini sites (P2-P2'), showing significantly enriched (top) or depleted (bottom) amino acids (p ≤ 0.05). Proteoforms resulting from N-terminal methionine excision were excluded prior to the analysis. (B) BP14 signal peptide prediction by Phobius, highlighting initiator and alternative initiator methionine in yellow. A B MVDGIGLFVMQFFSLSHAKGDGVMATKSIIKDVNIRDNKLCRTFASAIENASGRRGKDVQLSRSFKEISGEKIRELFGDKA 020 40 60 81 0 Signal peptide Identified proteoform 0.5 1.0 Signal peptide Non-cytoplasmic P2 P1 P1'P2' % Difference 30 15 -15 -30 ∑ Analyzed sites = 25 V | PROTEOGENOMIC ANALYSIS OF B. PRODUCTA P A G E | 114 FIGURE V-7 | Potential Alternative Initiation in SEP Proteoforms. Proteoforms for (A) BP4 and (B) BP7. For each proteoform, b-ions and y-ions are displayed in blue, while c-ions and z-ions are displayed in red, along with corresponding E-value, P-Score, and residue cleavage information. Overall, multiple SEP proteoforms were identified, some potentially featuring N-terminal formylation, acetylation, and disulfide bonds (FIGURE V-4). Specifically, proteoforms of BP3, BP4, BP7, BP10, and BP12 exhibited N-terminal formylation, while a proteoform of BP3 showed N-terminal acetylation. In addition to its role in protein synthesis, the N-formyl group may also serve as an indicator for cotranslational membrane insertion, with the N-terminal formyl group frequently retained (Bienvenut et al., 2015). Disulfide bridges were detected for proteoforms of BP12, BP24, and BP46, but the precise assignment of these linkages to specific cysteine residues in BP12 was not feasible with the available data (FIGURE A-21A-B). Secondary structure prediction using Alphafold revealed the presence of α-helix and β-sheet formations, as well as potential disulfide bonds between Cys35-Cys51 and Cys38-Cys54 for BP12 (FIGURE V-8 and FIGURE A-21C-E). FIGURE V-8 | Alphafold Structure Prediction for BP12. The predicted disulfide bridges between Cys35Cys51 and Cys38-Cys54 are indicated by lines connecting the orange-colored cysteine residues. A Canonical (E-value: 3.7E-21, P-Score: 5.4E-117, Residue cleavage: 87%) Formylated canonical (E-value: 6.5E-20, P-Score: 1.7E-89, Residue cleavage: 67%) Canonical (E-value: 8.7E-22, P-Score: 2.6E-135, Residue cleavage: 90%) Formylated canonical (E-value: 1.9E-22, P-Score: 6.4E-81, Residue cleavage: 57%) B Alternative initiation - Met5 (E-value: 3.3E-20, P-Score: 1.9E-95, Residue cleavage: 71%) Missing Alternative initiation - Met5 (E-value: 3.2E-21, P-Score: 1.4E-118, Residue cleavage: 81%) Missing C N Loop 1 Loop 2 1 MGIKVKVNFD KRKLESAIKD QARESLRNRS YDAKC35PFC38HT TFSAHPGPNVC51 PHC54RKTVDL NLNIKL 66 PROTEOGENOMIC ANALYSIS OF B. PRODUCTA | V P A G E | 115 3.3 Impact of Cultivation Conditions on SEP Translation Among the 44 SEP identified by BUP, 11 were consistently detectable across all seven growth conditions (FIGURE V-9A). Five of these SEP were exclusively identified under acidic growth conditions, while seven SEPs were consistently present across all conditions except for the acidic YCFA condition (pH 6.0 with the presence of SCFAs). Additionally, specific SEP were exclusively identified in the presence of particular components, such as yeast extract or LPS. These observations point towards the significant impact of environmental variables on SEP production dynamics. However, not all SEPs were influenced by these external factors. For example, three SEP (BP26, BP36, and BP42) were consistently produced in the YCFA medium (pH 7.0), regardless of the presence or absence of SCFAs, suggesting their independence from these specific factors. Additionally, all previously described SEP except BP15 were identified, including five SEPs (BP3, BP5, BP8, BP11, and BP12) previously detected only in co-culture with other bacteria (SIHUMIx) (Petruschke et al., 2021) (FIGURE V-9B). Their identification in single bacterial cultivation data suggests that their biosynthesis may not be solely reliant on interspecies interactions or communication within the microbiome. Furthermore, it suggests that the production of multiple SEP might have been influenced by specific bacterial growth and stress factors (FIGURE V-9B). For instance, BP11 could be linked to the presence of yeast extract in the medium at pH 7.0, while BP5 was only detected at pH 6.0 (FIGURE V-9B), indicating a potential involvement in the acid stress response. FIGURE V-9 | Presence of SEP across Diverse Culture Conditions. Bottom-up identifications across the seven different culture conditions of (A) all detected SEP (44 in total) and (B) previously described SEP (14 in total). Highlighted SEP are reported to be exclusively produced within the microbiome community. Set size indicates the total number of SEP identified in each condition, with intersection size representing shared SEP across different conditions. With permission from (Genth et al., 2023). 11 7 7 5 33221111 0 5 10 15 BP3, BP8, BP12 21BP11 1BP5 0 5 10 e a b d f c g e a b g f d c Intersection size Intersection size 17 30 19 30 34 34 37 10 12 10 11 12 12 13 A B Set size Set size BHI medium (pH 7.0 +SCFAs) a BHI medium (pH 7.0 -SCFAs) b BHI medium (pH 7.0 +Yeast) c BHI medium (pH 7.0 +LPS) d YCFA medium (pH 6.0 +SCFAs) e YCFA medium (pH 7.0 +SCFAs) f YCFA medium (pH 7.0 -SCFAs) g V | PROTEOGENOMIC ANALYSIS OF B. PRODUCTA P A G E | 116 3.4 Biochemical Predictions of Identified SEP The potential functions of the identified SEP were analyzed by predicting physicochemical properties and by examining sequence homologies based on reference domains, families, specific sites, or motifs. A distinct bimodal distribution of isoelectric point (pI) values was observed, indicating a notable shift of the SEP towards more basic pI values compared to the B. producta reference proteome (FIGURE V-10A). Furthermore, examination of the grand average of hydropathy (GRAVY) scores revealed a broader spectrum of hydrophilic properties (FIGURE V-10B), which aligned with the amino acid composition of the SEP characterized by a reduced frequency of hydrophobic amino acids like alanine, leucine, and isoleucine, alongside a higher occurrence of lysine and arginine residues (FIGURE V-10C). One particularly noteworthy finding was the potential antimicrobial peptide (AMP) activity exhibited by 25 SEP, as predicted by AMPfun (Chung et al., 2020) or AMP scanner v.2 (Veltri et al., 2018) (FIGURE V10D). Among these, BP16, BP34, and BP40 exhibited consistent AMP predictions (FIGURE V10A). Notably, BP40 also exhibited the potential to target both Gram-positive and Gramnegative bacteria. Additionally, approximately 71% of the identified SEP were predicted to be non-cytoplasmic (FIGURE V-10E). Furthermore, three non-cytoplasmic SEP (BP14, BP37, BP45) were predicted to have signal peptides, hinting at their potential involvement in cell-cell and cell-host communication (Hayes et al., 2010). FIGURE V-10 | Biochemical Predictions of Identified SEP. Distribution of (A) isoelectric points (pI), (B) grand average of hydropathy (GRAVY), and (C) amino acid composition between the B. producta reference proteome and the identified SEP. (D) Overlap of antimicrobial peptide (AMP) predictions. (E) Distribution of cellular localizations by Phobius. 71.1% 64.4% 6.7% 2.2% 26.7% Signal peptide prediction Cytoplasmic Non-cytoplasmic Transmembrane Signal peptide Non-signal Peptide D E 14 (31%) 3 (7%) 8 (17%) AMP Scanner 11 SEP AMPfun 17 SEP B A C 0% 2% 4% 6% 8% 10% G I L V A P F H W Y C M K R D E N Q S T 2 4 6 8 10 12 14 0 50 100 Distribution (%) pI -2 -1 0 1 2 0 50 100 Distribution (%) GRAVY Reference proteome Identified SEP PROTEOGENOMIC ANALYSIS OF B. PRODUCTA | V P A G E | 117 4 Discussion and Conclusion 4.1 Challenges and Advances in SEP Identification The presence of unannotated protein-coding sORFs were analyzed by validating their respective translated SEP products at the protein level. Utilizing various growth conditions, alongside the application of both BUP and TDP methodologies, a total of 45 SEP were successfully identified (FIGURE V-4). While BUP predominantly detected the majority of SEP primarily due to its higher sensitivity compared to TDP (Cassidy et al., 2023), TDP exclusively identified one SEP (BP46) (FIGURE V-2). Among the 45 SEP identified, 31 SEP (BP16-BP46) were identified for the first time, while 14 SEP (BP1-BP14) had been previously described (Petruschke et al., 2021). Both BUP and TDP approaches complemented each other in the identification of SEP, indicating the strengths of each methodology (FIGURE V-2). To ensure high-confidence identification of SEP, strict filtering criteria were implemented, aligning with recent recommendations in the field (Chen et al., 2023). Despite 17 out of 44 BUP identifications relying on a single peptide to support the protein-coding potential of the sORFs, these peptides exhibited an average of 69 PSMs. Although single-peptide identifications can be susceptible to increased false discovery rates (Nesvizhskii, 2010; Hadjeras et al., 2023), the majority of MS-based SEP identifications typically rely on a single unique peptide (Slavoff et al., 2013; Cassidy et al., 2019). This is due to the small protein size of SEP, resulting in a limited generation of tryptic peptides suitable for identification. To address this limitation, the use of multiple proteases (Bartel et al., 2020; Kaulich et al., 2021) or the integration of both BUP and TDP (Cassidy et al., 2016, 2021b) has proven effective in significantly improving sequence coverage and identification confidence of SEP. Incorporating TDP not only strengthens the evidence for the existence of SEP but also enables the detection of Nand Cterminal neo-termini and PTMs, achievements often challenging with BUP alone (Tholey and Becker, 2017). Application of TDP in this study identified 26 canonical proteoforms (FIGURE V-4) and increased the average sequence coverage by 19% compared to BUP. Furthermore, the identification of truncated proteoforms, potentially originating from proteolytic processing or alternative initiation, may suggest the potential impact of structural variations on the biological function of SEP (Melo et al., 2023). The identification of several proteoforms with PTMs suggests that the identified SEP can undergo posttranslational protein processing. The results of this study, alongside those of Petruschke and colleagues (Petruschke et al., 2021), suggest that the production of certain SEP is not solely dependent on interspecies interactions within the SIHUMIx co-culture system. Instead, specific SEP can also be produced under monoculture conditions and are influenced by various growth conditions and extracellular factors. Factors such as the presence of endotoxins (LPS) or pH variations, V | PROTEOGENOMIC ANALYSIS OF B. PRODUCTA P A G E | 118 known to play significant roles in gut inflammatory diseases such as IBD (Bai et al., 2016; Candelli et al., 2021), appear to act as stimuli or regulators for the production of specific SEP (FIGURE V-9). Moreover, a high number of identifications independent of culture conditions indicate a universal role for SEP in the survival and growth of B. producta. Overall, the results from this study, alongside those of Petruschke and colleagues, suggest a complex interplay between SEP and environmental factors (Petruschke et al., 2021). Nonetheless, it's worth considering why some of these potentially universal SEP were not identified in Petruschke and colleagues' study (Petruschke et al., 2021). Various factors may have limited the depth and detection capabilities of their proteomic analysis. One critical factor impacting proteomic workflows is the database search. Integrating a large proteogenomic database like SIHUMIx, including eight bacterial strains of interest, poses challenges in proteogenomic studies and subsequent FDR analysis (Nesvizhskii, 2014). This challenge arises due to the increased number of candidates competing for matching to an experimental MS/MS spectrum, increasing the risk of incorrect matches and distinguishing true from false identifications (Nesvizhskii, 2010). Consequently, large proteogenomic database searches may generate fewer non-canonical and total peptides compared to conventional searches with a reference database (Aggarwal et al., 2022). It's also essential to note that the absence of certain SEP in specific culture conditions does not necessarily imply their complete absence from those conditions. Instead, their presence may be at concentrations falling below the limits of detection of the methods or instruments used for analysis. Furthermore, interfering factors such as co-eluting compounds that compete for ionization can lead to reduced ionization efficiency for the target peptides (Keller et al., 2008), resulting in reduced signal intensities and sensitivity for low-abundance peptides, making identification challenging (Cassidy et al., 2023). To address these challenges, various techniques can be applied such as isolating, enriching, or depleting proteins of interest (Cassidy et al., 2019), applying a second chromatographic separation (Cassidy et al., 2021a), or using gas-phase separation (Swearingen and Moritz, 2012). These techniques can extend the limits of detection for low-abundant species in complex samples. To address some of the challenges Petruschke and colleagues conducted a comprehensive investigation of various proteomic techniques and approaches. They compared different enrichment strategies (C8-cartridge and GELFrEE), global proteomics methods (SP3, FASP, in-gel, and in-solution), and multiple protease cleavage approaches (trypsin and Asp-N) to identify small proteins in the SIHUMIX system (Petruschke et al., 2020). Their investigation provided valuable insights into the strengths and limitations of these techniques. By applying them to the SIHUMIX system, they identified several novel SEP, thereby contributing to the field of human gut microbiome research (Petruschke et al., 2021). They also acknowledged that differences in growth and cultivation conditions could impact the detection of specific SEP. PROTEOGENOMIC ANALYSIS OF B. PRODUCTA | V P A G E | 119 The results presented in this study complement the work of Petruschke and colleagues, suggesting that the production of certain SEP in B. producta is not exclusively reliant on interspecies interactions. Environmental factors can also impact the production of certain SEP in B. producta. Although SEP research is complex, and the challenges in methodology are acknowledged, both studies provide unique insights into the intricate nature of SEP biology. Therefore, it is crucial to regard these efforts as complementary, with each study contributing significantly to our understanding of SEP production. A comprehensive approach considering various experimental factors, including growth conditions, sample preparation techniques, and proteomic workflows, will be pivotal in advancing the field and uncovering the full spectrum of SEP. 4.2 Future Directions of SEP Research Given the focus of this study on soluble proteins, it is not surprising that the biochemical properties of the identified SEP exhibited a decreased relative frequency of hydrophobic amino acids and an increased frequency of charged amino acids, particularly lysine and arginine residues (FIGURE V-10C). These charged residues, especially those located at the C-termini of peptides resulting from tryptic digestion, typically enhance proton affinity, ionization efficiency, and CID fragmentation, leading to improved MS sensitivity (Dupree et al., 2020). The small overlap of AMP predictions for three SEP (BP16, BP34, BP40; FIGURE V-10A), can be likely attributed to variations in training sets, physicochemical properties, and amino acid frequencies at each sequence position, thus influencing the prediction model for classifying antibacterial peptides (Gabere and Noble, 2017). Notably, one SEP (BP27) was predicted to be a short transmembrane protein (FIGURE V-10E). Short membrane-associated SEP are predicted to constitute 35% of the SEP population in the human microbiome (Sberro et al., 2019). Given their potential roles in essential cellular processes such as transport, signaling pathways, cell division, respiration, sporulation, and membrane integrity (Yadavalli and Yuan, 2022), these transmembrane SEP should be targeted for future analysis. Consequently, specialized, MS-compatible membrane protein enrichment protocols should be prioritized in future studies to analyze these membraneassociated SEP (Capri and Whitelegge, 2017; Ahrens et al., 2022; Meier-Credo et al., 2022). Despite their potential significance, one of the major challenges associated with SEP lies in their lack of sequence homology with known proteins. Their small size and compact folding make traditional annotation and structure prediction tools less effective, as these tools typically depend on homology searches (Ahrens et al., 2022). This limits the understanding of the SEP function and emphasizes the need for experimental validation. Functional proteomics emerges as a promising approach to address these challenges and reveal the functional properties, biological roles, and mechanistic contributions of SEP. V | PROTEOGENOMIC ANALYSIS OF B. PRODUCTA P A G E | 120 Various proteomic methodologies, such as biochemical fractionation of soluble complexes coupled with mass spectrometry (Havugimana et al., 2022), affinity purification mass spectrometry (AP-MS) (Gnanasekaran and Pappu, 2023), or cross-linking mass spectrometry (XL-MS) (Piersimoni et al., 2022), are available to detect native protein complexes within cellular extracts and generate networks of protein-protein interactions (Low et al., 2021). These methodologies can significantly contribute to the identification of SEP within established protein networks, thereby enhancing our understanding of their roles and contributions to cellular functions (Garcia-del Rio et al., 2023; Leblanc et al., 2023). Additionally, genetic systems like the yeast two-hybrid assay can complement these efforts by mapping proteinprotein interactions, providing insights into the roles of SEP within the cellular interactome (Mehla et al., 2015). Furthermore, integrating transcriptomic data created using RNA-Seq or ribosome profiling with proteomic analysis of the same biological samples can increase the number of experimentally validated SEP (Guilloy et al., 2023; Hadjeras et al., 2023). This integrated omics approach enables a more comprehensive exploration of the SEP proteome, leading to a better understanding of their biological functions. Future research in this field will likely provide more insight into the significance of these small proteins. OUTER MEMBRANE VESICLES ANALYSIS | VI P A G E | 127 comprised 10 mM DTT and 50 mM IAA within a 100 mM TEAB buffer environment. All samples were digested with trypsin at a 1:40 enzyme-to-substrate ratio (20 h, 37°C, 800 rpm), with the addition of 0.01% (w/v) n-dodecyl-β-D-maltoside (DDM) or 0.5% (w/v) SDC. For SDC digests, samples were additionally processed using a modified phase transfer protocol (Masuda et al., 2008) (chapter II.3.3). TABLE VI-1 | Overview of Protocol-specific Sample Processing Steps. All listed values represent end concentrations for the in-solution digestion (ISD), single-pot, solid-phase-enhanced protein extraction (SP3), or filter-aided sample preparation (FASP) protocol. PROTOCOL STEP CONDITIONS ISDCLASSIC ISD-I MPROVED SP3SDS - TEAB SP3SDS - SDC SP3SDS - DDM SP3SDC FASPSDC FASP-SDS-SDC OMV LYSIS Freeze thawing ü 2% SDC ü ü ü 1% SDS ü ü ü ü RED. & ALK. 10 mM DTT & 50 mM IAA ü ü ü ü ü ü ü ü SDS REMOVAL 8 M urea, 100 mM TEAB ü TRYPTIC DIGESTION 0.5% SDC, 100 mM TEAB ü ü ü ü ü 0.001% DDM, 100 mM TEAB ü 100 mM TEAB ü ü SDC REMOVAL 0.5% TFA & 100% ethyl acetate ü ü ü ü ü 2.5 Caco-2 Wound-Healing Assay An important feature of OMVs is that the proteins associated with them exhibit various biological activities. Particularly, OMVs produced by E. coli can exhibit an inhibitory effect on cell proliferation and induce pro-inflammatory responses in intestinal epithelial cells (Cañas et al., 2016; Patten et al., 2017). To evaluate the potential biological activity of isolated E. coli OMVs, a wound-healing assay utilizing the human intestinal epithelial cell line Caco-2 was conducted. Wound closure was measured immediately after removing the insert, as well as at 12 and 30 hours after incubation with increasing OMV concentrations (10, 50, and 100 µg/ml). Positive and negative controls included incubation with 5 ng/ml TGFβ and 1 µg/ml LPS, respectively. Cells solely incubated with the medium (0.1% FCS) and the OMV elution buffer served as references for comparison with the treated groups. These controls facilitated the calculation of relative wound closure. Further details on the cultivation of Caco-2 cells and the wound healing assay can be found in chapter II.2.3. VI | OUTER MEMBRANE VESICLES ANALYSIS P A G E | 128 3 Results 3.1 Proteomic Workflow Compatibility Given that the reagents of the OMV kit are undisclosed, conducting a thorough compatibility check, specifically focusing on LC-MS/MS analysis and protein digestion, was essential for its integration into a proteomic workflow. MALDI MS analysis of diluted eluate from a blank kit run revealed no characteristic PEG ion series or other polymeric impurities. While the absence of characteristic impurities indicates compatibility of the OMV kit with mass spectrometry techniques (Keller et al., 2008), it does not guarantee complete compatibility. Additional LCMS analysis of two dilutions (1:10 and 1:100) using cytochrome C digest for retention time and HeLa digest for peptide identification controls showed no peptide retention time shifts and comparable peptide identification. This indicates that the elution behavior of peptides and peptide detection remained largely unaffected, at least for subsequent LC-MS runs. Monitoring signal stability and background noise revealed a singly charged peak (309.125 m/z) at 53 minutes, with intensities of approximately 1.8 x 108 and 6.4 x 108 in the 1:100 and 1:10 dilutions, respectively. The observed minimal impact on chromatographic runs and MS analysis confirms the compatibility of the OMV kit with LC-MS analysis. Furthermore, the compatibility of the OMV elution buffer with tryptic digestion was analyzed using protein mixtures consisting of six proteins (6P), and four proteins (4P), the latter excluding myoglobin and alcohol dehydrogenase (FIGURE VI-2A). The mixtures were digested using either TEAB-buffered OMV elution buffer (KIT) or 100 mM TEAB (TEAB) and were evaluated through SDS-PAGE analysis (FIGURE VI-2B). While the comparison of the two different digestion conditions showed slightly reduced protein band intensities with the TEABbuffered OMV elution buffer (FIGURE VI-2B), the results still confirmed the compatibility of the OMV kit with tryptic digestion, allowing integration of the kit into a proteomic workflow. FIGURE VI-2 | SDS-PAGE Analysis of Protein Digestion using the Exobacteria Elution Buffer. (A) Molecular weight distribution and composition of the six-protein mixture. (B) Evaluation of the tryptic digestion of the 6-protein (6P) or 4-protein (4P) mixtures (excluding myoglobin and alcohol dehydrogenase) using TEAB-buffered OMV elution buffer (KIT) or 100 mM TEAB. kDa 212 118 66 A 20 BSA ADH CA & β-Cas CA & Myo CytC 43 14 29 TEAB KIT B OUTER MEMBRANE VESICLES ANALYSIS | VI P A G E | 129 3.2 Biophysical Characteristics and Kit Loading Capacity Biophysical Characteristics – The isolated OMVs were first plated on LB agar to confirm the absence of bacterial contamination. Subsequently, they were characterized using dynamic light scattering with NTA, which allowed direct, real-time visualization of the isolated OMVs. FIGURE VI-3A illustrates a screenshot from one of the recorded videos. Examination of the isolated OMVs revealed nanoparticles ranging from 117.2 ± 2.1 nm (mode ± standard error) for OMVs isolated from acetate supernatant to 95.3 ± 4.1 nm for those isolated from glucose supernatant (FIGURE VI-3B and TABLE VI-2). The average particle concentration was higher for OMVs isolated from supernatant obtained from acetate cultivation (9.4 x 1010 ± 0.9 particles/ml) compared to those isolated from supernatant obtained from glucose cultivation (7.8 x 108 ± 0.8 particles/ml). The observed differences in particle concentration between the two supernatants, while providing a general reference, require additional replicates for thorough validation. Furthermore, NTA measurements indicate a relatively narrow size distribution of particles, suggesting a more monodisperse sample rather than a polydisperse one. FIGURE VI-3 | Nanoparticle Tracking Analysis of E. coli OMVs. (A) Representative visualization of OMV particles captured from the recorded video. (B) Particle size distribution of OMVs isolated from E. coli supernatant cultured in acetate or glucose M9 medium. TABLE VI-2 | Summary of Nanoparticle Tracking Analysis PARAMETER ACETATE GLUCOSE Mode particle size 117.2 ± 2.1 nm 95.3 ± 4.1 nm D10 82.9 ± 3.4 nm 79.2 ± 3.8 nm D50 124.3 ± 1.7 nm 103.4 ± 2.6 nm D90 204.3 ± 5.7 nm 207.4 ± 19.7 nm Particle concentration 9.4 x 1010 ± 0.9 particles/ml 7.8 x 108 ± 0.8 particles/ml A B 0 1 2 3 4 0 1 2 3 4 0200 400 600 800 1000 Size (nm) Particles x 10 7 Particles x 10 9 Acetate Glucose VI | OUTER MEMBRANE VESICLES ANALYSIS P A G E | 130 Loading Capacity – The assessment of the kit's loading capacity revealed that an increase in the initial supernatant volume from 20 to 40 or 60 ml resulted in a significant doubling of protein concentrations (FIGURE VI-4A) and particle concentrations (FIGURE VI-4B). Notably, the particle size and distribution of particles remained constant at 106 ± 3.0 nm (FIGURE VI-4C and FIGURE A-24), indicating that OMVs maintained their size and distribution even with an increased particle concentration. Furthermore, the purity ratio (the ratio of vesicle counts to protein concentration) of OMV preparations, which considers both protein contamination and loss of OMVs during isolation (Webber and Clayton, 2013), remained relatively constant at 4.3 x 106 ± 0.1 particles/µg of protein (FIGURE VI-4D). Despite variations in the loaded supernatant volume, a comparable particle size and purity ratio could be maintained, indicating the kit's efficiency in handling increased sample volumes. FIGURE VI-4 | Loading Capacity Evaluation of the ExoBacteria OMV Isolation Kit. Different E. coli supernatant volumes (20, 40, and 60 ml) were processed with the OMV isolation kit and analyzed using NTA and BCA. (A) Protein concentration, (B) particle concentration, (C) particle mean size, and (D) normalized purity ratio (particle/protein ratio). Data represent mean ± standard error from 3 independent experiments. Significant differences were calculated via one-way ANOVA with Dunnett’s correction. * (p < 0.05); ** (p < 0.01). Biological Activities – The ability of isolated OMVs to exhibit biological activities was evaluated using an in vitro wound closure assay, which examined the migration and proliferation of human colonic Caco-2 cells. The evaluation of wound size reduction after 12 and 30 hours indicated a notable reduction of wound healing compared to untreated control cells (FIGURE VI-5A). The quantitative assessment of wound healing, with the medium arbitrarily set at 100% as a reference, revealed reduced wound closure with increasing concentrations of OMVs. At concentrations of 10 µg/ml, 50 µg/ml, and 100 µg/ml, OMVs exhibited decreasing percentages of wound closure after 12 hours (45 ± 1.1%, 68 ± 3.0%, and 82 ± 2.8%, respectively FIGURE VI-5B). Despite statistical analysis using one-way ANOVA with Dunnett’s correction, no significant changes in wound healing (p < 0.05) were observed. This lack of statistical significance may be attributed to the limited sample size of only three biological experiments. Although the results were not statistically significant, they suggest a possible dose-dependent inhibitory effect of OMVs on wound closure, with higher concentrations having a greater effect. While the OMV elution buffer showed a modest 64 113 153 0 50 100 150 200 20 mL 40 mL 60 mL Protein concentration (µg/ml) 4.2 4.2 4.4 0 2 4 6 8 20 mL 40 mL 60 mL Purity ratio (10 6 particles/µg protein) 20 ml 40 ml 60 ml 109 106 102 0 50 100 150 200 20 mL 40 mL 60 mL Particle mean size (nm) 2.7 4.8 6.6 0 2 4 6 8 20 mL 40 mL 60 mL Particle concentration (10 8 /ml) A B CD * * ** OUTER MEMBRANE VESICLES ANALYSIS | VI P A G E | 131 stimulatory effect on wound closure after 30 hours (9 ± 2.0% FIGURE VI-5B), this effect may be attributed to common buffer components like calcium and phosphoric acid, which can influence the wound healing process (Navarro-Requena et al., 2018; Sim et al., 2022). The negative control (LPS) resulted in a reduction of 22 ± 2.5%, while the positive control (TGFβ), known for its role in promoting proliferation (Penn et al., 2012), stimulated wound closure by 36 ± 6.2%. FIGURE VI-5 | Caco-2 Wound Healing Assay. (A) Time-lapse microscopy images of wound closure of Caco-2 cells with medium, TGFβ (5 ng/ml), LPS (1 µg/ml), and OMVs (50 µg/ml). Images were captured at 0 h, 12 h, and 30 h (10-fold magnification). (B) Wound closure after 12 and 30 h incubation with varying OMV concentrations (10, 50, and 100 µg/µl), LPS (1 µg/ml), OMV elution buffer, medium alone, or TGFβ (5 ng/ml). Medium alone was arbitrarily assigned as 100%. Data represent mean ± standard error from three independent experiments. Bars represent 200 μm. 3.3 Proteomic Analysis of E. coli OMVs The proteomic changes in E. coli OMV protein content under different growth states and culture conditions were analyzed by a bottom-up proteomic analysis. Changes in protein abundance and localization were estimated by annotating subcellular locations using STEPdb 2.0 (Loos et al., 2019), which categorizes proteins into 13 distinct subcellular classes (FIGURE A-23A). Due to the dynamic nature of protein localization and their ability to move to various extracytoplasmic compartments, several proteins may exhibit multiple subcellular locations (FIGURE A-23B). To facilitate annotation, proteins from different subcellular compartments were classified into four distinct subcellular topological groups: the cytoplasm (F1, A, R, and N), the inner membrane (B), the periplasm (I, G, F2, F3, and E), and the outer membrane/extracellular group (H, X, and F4) (FIGURE A-23C). These subcellular topological classifications considered most of the dynamic protein movement, providing robust insights into their cellular localization. For mid-logarithmic growth phases, a total of 334 and 298 proteins were identified in OMVs isolated from supernatant obtained from glucose and acetate cultivation, respectively (FIGURE VI-6A). Among these proteins, 258 (69%) were identified in both isolations, while 76 proteins were identified exclusively under glucose and 40 proteins exclusively under acetate conditions A B 0% 20% 40% 60% 80% 100% 120% 140% 160% 12 h 30 h % Wound healing OMV 100 (µg/ml) OMV 50 (µg/ml) OMV 10 (µg/ml) LPS Medium OMV elution buffer TGFβ LPSOMV MediumTGFβ 0 h 12 h 30 h VI | OUTER MEMBRANE VESICLES ANALYSIS P A G E | 132 (FIGURE VI-6A). The relevance of the exclusively identified OMV proteins remains uncertain due to the absence of clear functional or pathway enrichments. Comparison of the subcellular locations of identified OMV proteins showed that OMVs isolated from supernatant obtained from acetate or glucose cultivation were enriched in periplasmic and cytoplasmic proteins, whereas only 9% of the total proteins were classified as outer membrane or extracellular proteins (FIGURE VI-6B-C). The majority of identified cytoplasmic proteins were associated with ribosomal functions or involved in glycolysis, such as enolase (P0A6P9), glyceraldehyde-3-phosphate dehydrogenase (P0A9B2), and phosphoglycerate kinase (P0A799). However, most other abundant and essential cytoplasmic proteins required for bacterial survival were not detected (Goodall et al., 2018). FIGURE VI-6 | Subcellular Topological Distribution of OMV Proteins based on Total Protein Identifications. (A) Overlap of proteins identified in OMV isolated at mid-logarithmic growth from glucose and acetate cultivation. (B) Subcellular topological distribution of OMVs isolated at midlogarithmic growth from glucose and (C) acetate cultivation Determination of median abundance values from the two independent OMV isolates, each measured in three technical replicates, allowed evaluation of protein abundance within subcellular topological categories. Comparing the distribution based on the median abundance with the distribution based on the total number of identified proteins of each subcellular topological category revealed a distinct pattern (FIGURE VI-7). For OMVs isolated from supernatant obtained during the mid-logarithmic growth phase of glucose cultivation, 68.4% of the median abundance originated from the outer membrane/extracellular, 20.5% from periplasmic, 4.3% from the inner membrane, and 6.8% from cytoplasmic proteins (FIGURE VI7A). Similarly, OMVs isolated from supernatant obtained during acetate cultivation exhibited a distribution where 73.5% of the median abundance originated from outer membrane/extracellular, 18.8% from periplasmic, 2.8% from inner membrane, and 4.9% from cytoplasmic proteins (FIGURE VI-7B). This observation suggests that, while a substantial number of cytoplasmic proteins were present within the OMVs isolated from supernatant obtained from glucose (143 proteins) or acetate cultivation (119 proteins), their median abundance is notably lower compared to that of the 29 glucose or 27 acetate outer membrane/extracellular proteins. However, it is essential to note that solely relying on median 39.3% 3% 48.8% 8.9% C A Cytoplasm Inner membrane Periplasm Outer membrane 8.6% 46.3% 2.9% 42.2% B Acetate 298 Glucose 334 40 (11%) 258 (69%) 76 (20%) OUTER MEMBRANE VESICLES ANALYSIS | VI P A G E | 133 abundance may underestimate the significance of less abundant proteins which could also play important functional roles or contribute to biological processes. Therefore, both distributions of the subcellular topological categories should be considered complementary, providing distinct insights into protein abundance and diversity. Despite an increase in the overall protein content of the OMV isolates (from 0.23 to 0.83 µg/µl), the subcellular topological distribution of OMVs isolated from supernatant obtained from glucose cultivation remained stable throughout the mid-logarithmic to stationary phase (FIGURE VI-7A). As cultivation time increased, only a 3.2% increase in the median abundance of cytoplasmic proteins and a 4.8% decrease in outer membrane/extracellular proteins were observed for OMVs isolated from supernatant obtained from glucose cultivation (FIGURE VI7A). In contrast, OMVs isolated from supernatant obtained from acetate cultivation exhibited a notable 13.5% decrease in the median abundance of outer membrane/extracellular proteins, accompanied by a 6.4% increase in cytoplasmic proteins and a 9.4% increase in inner membrane proteins spanning the mid-logarithmic to death phase (FIGURE VI-7B). The decrease in outer membrane/extracellular proteins may be attributed to the direct growth inhibition that occurred upon entering the stationary phase (FIGURE A-22). Additionally, the activation of programmed cell death mechanisms may have contributed to the increase in cytoplasmic and inner membrane proteins (Juodeikis and Carding, 2022). FIGURE VI-7 | Subcellular Topological Distribution of OMV Proteins based on Median Abundance. (A) OMVs isolated from supernatant obtained during the mid-logarithmic, pre-stationary, and stationary growth phases of glucose cultivation. (B) OMVs isolated from supernatant obtained during the midlogarithmic, late-stationary, and death growth phases of acetate cultivation. Cytoplasm Inner membrane Periplasm Outer membrane 68.4% 20.5% 4.3% 6.8% A 8.3% 2.2% 15.6% 73.9% Pre-stationaryMid-logarithmic 10%2.5% 23.5% 64% Stationary 4.9% 2.8% 18.8% 73.5% BMid-logarithmic 9.5% 2% 20.4% 68.1% Late-stationary 11.3% 12.2% 16.4% 60% Death phase VI | OUTER MEMBRANE VESICLES ANALYSIS P A G E | 134 Examination of major components of outer membranes (TABLE A-10), suggested as ubiquitous markers for OMV validation (Daleke-Schermerhorn et al., 2014; Hong et al., 2019), revealed a consistent increase in iBAQ values in OMVs isolated from supernatant obtained from glucose cultivation (FIGURE VI-8A). Certain marker proteins, such as OmpT and OmpF, exhibited increased iBAQ values in isolated OMV samples (FIGURE VI-8A) compared to full proteome preparations of E. coli (FIGURE VI-8B), which may suggest the selective sorting and encapsulation of certain proteins into OMVs. Interestingly, other marker proteins, such as OmpA, OmpX, and OmpC, consistently showed high iBAQ values for both proteomes. FIGURE VI-8 | iBAQ Distribution of Potential E. coli OMV Protein Markers. (A) The OMV proteome isolated from the supernatant obtained from glucose cultivation and (B) the complete E. coli proteome. Protein names are listed in TABLE A-10. 3.4 Improving OMV Lysis and Trypsin-Based Digestion The impact of the non-ionic detergent SDC on OMV integrity and the susceptibility of OMV proteins to enzymatic degradation were evaluated (FIGURE VI-1). OMV samples subjected to tryptic digestion exhibited protein bands within the 43 to 66 kDa range (FIGURE VI-9A). The application of 2% (w/v) SDC before enzymatic digestion resulted in the reduction of these bands. A control experiment involving only 2% SDC confirmed that the observed bands corresponded to proteins encapsulated within the OMVs. Protein bands with a molecular mass of approximately 23.5 kDa correspond to trypsin itself. Overall, the results indicate the capability of SDC to induce OMV lysis and the subsequent susceptibility of proteins initially protected within OMVs to enzymatic degradation. Evaluation of SDC and DDM impact on tryptic digestion of the six-protein mixture, compared against a TEAB control, indicated improved tryptic digestion using 0.5% (w/v) SDC (FIGURE VI9B). Conversely, 0.01% (w/v) DDM exhibited protein bands with similar intensities to those of the TEAB control, suggesting minor differences in protein abundance between the two digestion conditions (FIGURE VI-9B). Especially for TEAB-buffered OMV elution buffer (KIT), 0100 200 300 400 500 2 4 6 8 10 0500 1000 1500 2 4 6 8 10 OmpA OmpC Lpp OmpT OmpF SlyBSlp OmpW LpoA Tsx RlpA BamD BamB RcsF BamE BamC BamA LolB OmpX OmpA OmpC OmpX Slp Lpp SlyB OmpT BamC BamB BamA BamE RcsF Tsx BamD OmpF LolB LpoA RlpA OMV log10 iBAQ Ranked proteins A Porines Lipoproteins Assembly proteins Full proteome log10 iBAQ Ranked proteins B OUTER MEMBRANE VESICLES ANALYSIS | VI P A G E | 135 the results indicate an improved trypsin digestion efficiency in the presence of SDC compared to DDM or TEAB. FIGURE VI-9 | Evaluation of Non-Ionic Detergents on OMV Lysis and Tryptic Digestion. (A) SDC-Mediated OMV Lysis. From left to right: molecular weight ladder (in kDa), untreated OMV sample, tryptic-treated OMV sample, OMV sample treated with 2% (w/v) SDC followed by tryptic treatment, and OMV sample treated with 2% (w/v) SDC. (B) Non-ionic detergents effect on tryptic digestion using TEAB-buffered OMV elution buffer (KIT) or 100 mM TEAB (TEAB). To improve trypsin-based digestion of OMV samples, various sample preparation protocols were tested, including on-bead SP3, on-membrane FASP, and in-solution methods. These protocols utilized both SDC and DDM for effective lysis and tryptic digestion of OMVs. Total identified protein groups ranged from 347 to 710 (FIGURE VI-10A), with peptide identifications ranging from 1140 to 5043 (FIGURE VI-10B). The ISD-improved protocol, employing SDC for both OMV lysis and tryptic digestion, exhibited the highest number of protein identifications (FIGURE VI-10A). The SP3 protocol, integrating SDS for OMV lysis and SDC for tryptic digestion (SP3-SDS-SDC), exhibited the most identified peptides (FIGURE VI-10B). Comparison of different SP3 protocols showed that the integration of SDC or DDM for tryptic digestion resulted in higher protein and peptide identifications compared to detergent-free digestion using only TEAB. (FIGURE VI-10A-B). Despite the potential interference from SDS, which could have affected trypsin's activity and led to incomplete protein digestion (Masuda et al., 2008), less than 3% of the identified peptides had 2 missed cleavages (FIGURE VI-10C). These results indicated efficient removal of SDS and successful protein digestion. An exception was the FASP-SDS-SDC method, which exhibited the lowest number of protein and peptide identifications, as well as the lowest sequence coverage (FIGURE VI-10D), across all three biological replicates. Small leftovers of SDS, which may not have been sufficiently removed, could potentially have interfered with tryptic digestion or led to signal suppression during LC-MS measurements (Rundlett and Armstrong, 1996; Masuda et al., 2008). A BSA ADH CA & β-Cas CA & Myo CytC 6P KIT TEAB B 118 29 43 20 66 kDa kDa 212 118 20 43 14 29 66 212 VI | OUTER MEMBRANE VESICLES ANALYSIS P A G E | 136 Furthermore, sample loss, due to small leftovers remaining in the filter reservoir or the filter itself, might have occurred during sample processing. FIGURE VI-10 | Peptide and Protein Identification Data of Different OMV Sample Preparations. (A) Protein identifications, (B) peptide identifications, (C) missed cleavage events, and (D) protein sequence coverage. Box-and-whisker plots capture lower quartile and upper quartile with the median displayed as a horizontal line and the mean depicted as a cross; whiskers represent minimum and maximum values that fall within 1.5 times the interquartile range. Overall, all methods achieved similar sequence coverages of 23% (FIGURE VI-10D), with each method covering largely similar fractions of the E. coli proteome (FIGURE A-25). Compared to the detergent-free ISD-classic protocol which employed freeze-thawing for OMV lysis, the integration of SDS or SDC for OMV lysis and SDC or DDM for digestion notably improved peptide and protein identifications, resulting in higher sequence coverage. For this reason, subsequent analysis focused on evaluating these improved methods. Notable methodological variations were observed in the subcellular topological distribution of OMV proteins, with cytoplasmic proteins ranging from 2.2% to 4.2%, inner membrane proteins from 0.2% to 27.8%, periplasmic proteins from 2.4% to 24%, and outer membrane proteins from 64% to 87% (FIGURE VI-11A). The ISD-improved method exhibited the highest abundance of outer membrane proteins (87%) and the lowest abundance of inner membrane proteins (0.2%) (FIGURE VI-11A). Interestingly, FASP-SDC and SP3-SDC, which also employed SDC for digestion and lysis, showed the highest abundance of inner membrane proteins at 26% and 457 710 471 597 597 531 657 347 3162 3906 3885 5043 4720 4466 4445 1140 ISD-Classic ISD-Improved SP3-SDS-TEAB SP3-SDS-SDC SP3-SDS-DDM SP3-SDC FASP-SDC FASP-SDS-SDC 0 200 400 600 800 1000 Σ Proteins A 0 1500 3000 4500 6000 Σ Peptides B C D 0 20 40 60 80 100 Sequence coverage (%) 0 20 40 60 80 100 Missed cleavage (%) 2 1 0