scieee AI-readable full text Open interactive document viewer

Protein Subcellular Localization Prediction: Historical, Current, and Prospective

M. Malkiya Ruth; B.N. Prathibha

Abstract

One of the biggest challenges of the post genomic era is the functional characterization of every single protein. Proteomics, the large-scale study of a cell’s proteins, aims to give these proteins accurate annotations on their interaction partners and roles in the cellular apparatus. Determining each protein’s sub cellular location is a crucial step in this process. Organelles, or sub cellular compartments, make up eukaryotic cells. A highly controlled and intricate cellular process is transport across the membrane into the organelles. Recently, there has been a lot of research in the field of computationally predicting sub cellular localization. Four key factors that are significant to the user—the computational approach, the fundamental biological reason, the localization coverage, and reliability— distinguish the publicly available prediction systems from one another. The key steps in the protein sorting process are briefly described in this review, along with a summary of the most often utilized techniques in this area

Full text

144 £UP® kupah fiy kw;Wk; mwptpay; (kfspu;) fy;Y}up> jpUney;Ntyp gz;ila tuyhWk; jkpou; gz;ghLk; Protein Subcellular Localization Prediction: Historical, Current, and Prospective M. Malkiya Ruth Research Scholar, Department of Computer Science Sadakathullah Appa College(Autonomous) Affliated to Monmaniam Sundaranar University, Tirunelveli, India Dr. B.N. Prathibha Assitant Professor, Department of Computer Science Sadakathullah Appa College (Autonomous), Tirunelveli, India Abstract One of the biggest challenges of the post genomic era is the functional characterization of every single protein. Proteomics, the large-scale study of a cell’s proteins, aims to give these proteins accurate annotations on their interaction partners and roles in the cellular apparatus. Determining each protein’s sub cellular location is a crucial step in this process. Organelles, or sub cellular compartments, make up eukaryotic cells. A highly controlled and intricate cellular process is transport across the membrane into the organelles. Recently, there has been a lot of research in the field of computationally predicting sub cellular localization. Four key factors that are significant to the user—the computational approach, the fundamental biological reason, the localization coverage, and reliability— distinguish the publicly available prediction systems from one another. The key steps in the protein sorting process are briefly described in this review, along with a summary of the most often utilized techniques in this area. Keywords: sub cellular localization, prediction methods Introduction A vast amount of sequence data has been produced by extensive global genomic and proteomic activities. One of the main forces behind molecular and computational biology has been the annotation of these sequences (1, 2). The goal of functional annotation initiatives is to clarify the possible functions of proteins in biological contexts, including metabolic pathways and interaction networks. Up to 10,000 distinct protein types can be produced by eukaryotic cells, and each one is intended for one or more pre-identified target organelles. The proper transport of a protein to its ultimate location is essential to its function because proteins have evolved to function best in a particular subcellular location. Protein targeting, also known as protein sorting, is the process of guiding a freshly generated protein to its intended organelle. A number of human illnesses, including cancer and Alzheimer’s disease, have been linked to failures in protein transport (3–5). Automated, high-throughput localization offers an enticing alternative to experimental methods. The creation of techniques for predicting subcellular localization has been a hot topic in recent years (6), and it has long been thought of as a bioinformatician’s detective job (7). For the eager inventors of prediction methods, the immense complexity of the protein sorting process, alternate transportation pathways, and the absence of complete data for every organelle pose significant hurdles. ©»º: 13 ]Ó¨¤uÌ: 2 ©õu®: ö\¨h®£º Á¸h®: 2025 P-ISSN: 2321-788X E-ISSN: 2582-0397 DOI: https://doi.org/10.5281/ zenodo.17311622 http://www.shanlaxjournals.com 145 £UP® Shanlax International Journal of Arts, Science and Humanities In this study, we outline the key steps involved in protein sorting, present a summary of the computational advancements made in this area, and conclude with some advice for prospective users. Background in Biology Eukaryotes have a minimum of ten primary subcellular localizations, some of which are further split into intraorganellar compartments. In contrast, bacteria are made up of a plasma membrane and a single intracellular compartment. The organelles are believed to have developed from ancestral bacterial endosymbionts of prokaryotic cells and serve unique, well-defined, and complementary roles in the cellular machinery.(8). The majority of the cell’s proteins are encoded in the nucleus DNA, only a tiny portion of which is encoded in mitochondrial and chloroplastic DNA. Each protein is guided to its ultimate destination by a complex and incredibly selective set of sorting and transportation processes (9, 10). The translation of mRNA into protein occurs in the cytoplasm, which envelops the nucleus. Cytoplasmic proteins can either stay in the cytoplasm, travel to other non-secretory pathway (nSP) organelles. George Palade (11), who was awarded the Nobel Prize for his work in 1974, conducted groundbreaking experimental experiments that identified the intracellular trafficking of noncytoplasmic proteins. The secretory pathway proteins are cotranslationally transported across the Endoplasmic Reticulum (ER) membrane and carry a targeting sequence in their precursor protein sequences. Unless they contain an ER-retention sequence, proteins in the ER are subsequently transported into the Golgi apparatus, plasma membrane, lysosome, vacuole, or extracellular space. Proteins are frequently transported by vesicular carriers, which have been demonstrated to move between the Golgi apparatus(9,10).If the nSP proteins have particular N-terminal targeting sequences for the mitochondria (mTP), lysosome, or chloroplast (cTP), they are transported from the cytoplasm post-translationally after being generated on free cytoplasmic ribosomes (14, 15). Once the protein has arrived at its location, certain signal sequence peptidases often break these targeting regions off the mature protein sequence (16–18). Transmembrane protein membrane insertion can be initiated by additional intrinsic sequences found in the mature protein, such as the hydrophobic stoptransfersequence. Chloroplasts have secondary targeting sequences that facilitate additional intraorganellar trafficking (19). It is necessary to import all nuclear proteins from the cytoplasm. The nuclear pore complex facilitates this import by identifying a nuclear localization signal (NLS), which is unique to nuclear proteins (20,21). The brief sequence of four to eight amino acids, known as the NLS, is often positively charged. It can be encoded as a single fragment (monopartite) or as two separate fragments (bipartite) (22,23). Although the NLS has been accurately identified in a number of nuclear proteins, certain nuclear proteins seem to have no NLS all. Additionally recruited from the cytoplasm, peroxisomal proteins have a brief C-terminal signal sequence that makes transport across the peroxisomal membrane easier(24,25). Moreover, glycosylation, one of the posttranslational modifications, is crucial for additional protein trafficking High specificity and evolutionary conservation are traits shared by all signal sequence(26)s. Conservation may be indirectly seen at the level of the amino acids’ metabolic characteristics rather than directly within the fundamental amino acid sequence. Certain targeted sequences exhibit a propensity to develop a secondary structure.. A few direct sequence motifs have been found, and the primary sequence plays a major role in the proper cleavage of the TPs. The proteins are exposed to particular biological circumstances by the organelles(27,28). Only mutations that are advantageous to the cell have been accepted throughout the evolution process. The fundamental theory that each protein has evolved over time to perform best in a particular subcellular location can be developed because it has been noted that proteins from different organelles differ in their general amino acid content (29). 146 £UP® kupah fiy kw;Wk; mwptpay; (kfspu;) fy;Y}up> jpUney;Ntyp gz;ila tuyhWk; jkpou; gz;ghLk; Computational Methods Four broad categories can be used to classify computational approaches for protein subcellular localization prediction: (i) prediction methods based on the total amino acid composition of the protein, (ii) known targeting sequences, (iii) sequence homology and/or motifs, and (iv) hybrid methods, which combine information from multiple sources from the first three categories(30). Nakashima and Nishikawa devised a technique for differentiating between intracellular and extracellular proteins, which was the first work to use the total amino acid composition for prediction. Cedano et al. introduced ProtLock, a tool for predicting five classes of subcellular localizations (extracellular, intracellular, integral membrane, anchored membrane, and nuclear based on the distance between the vectors representing the overall amino acid composition. Reinhardt and Hubbard introduced NNPSL, a method for predicting three prokaryotic (cytoplasmic, extracellular, and periplasmic) and four eukaryotic (cytoplasmic, extracellular, mitochondrial, and nuclear) subcellular localizations. The provided data set has been subjected to a number of different algorithms. Reinhardt and Hubbard, such as Markov chain models, Support Vector Machines, and Kohonen’s self-organizing maps. Sequence order effects have led to further advancements in the use of the overall amino acid composition. Chou et al. introduced an SVM-based technique that accounts for sequence order effects to predict twelve distinct subcellular localizations . The data set provided by Reinhardt and Hubbard has been subjected to a number of other techniques, such as Markov chain models , Support Vector Machines, and Kohonen’s self-organizing maps. Sequence order effects have led to further advancements in the use of the overall amino acid composition. Chou et al. introduced an SVM-based technique that accounts for sequence order effects to predict twelve distinct subcellular localizations. Park and Kanehisa recently proposed a similar strategy when they explained the PLOC method. Huang et al. used fuzzy k-NNs to characterize the dipeptide composition of the whole protein sequence for eleven distinct localizations. Based on the makeup of peptides of different lengths, the CELLO approach allows for the prediction of five subcellular localizations in Gram-negative bacteria: the cytoplasm, inner membrane. The structural information was originally added to the vectors representing the amino acid composition. Nuclear, extracellular, and cytoplasmic proteins were identified using the surface makeup of eukaryotic proteins. This method is justified by the fact that whereas surface residues of proteins have adapted to certain biochemical conditions, the inside of proteins have remained relatively stable throughout evolution. TargetP, the most complete approach based on Nterminal targeting sequences, enables the prediction of proteins such as secretory pathways, mitochondria, and chloroplasts. TargetP can be thought of as a combination of the SignalP, ChloroP, and SignalP techniques, all of which were introduced by the Gunnar von Heijne group. Predotar (http://www.inra.fr/predotar) and MitoProt, two techniques, both specifically distinguish between mitochondrial and chloroplast proteins. iPSORT, another technique in this field, provides TargetP-like localization category prediction. Protein sequence characteristics obtained from the AAindex database are used by the iPSORT to make predictions using knowledge-based rules. Marcotte et al. introduced a technique that uses protein phylogenetic patterns to determine the subcellular localization. Mott et al. predicted nuclear, secreted, and cytoplasmic proteins using SMART domains. Based on a set of nuclear localization sequences, the PredictNLS approach is specialized in identifying nuclear proteins. The Reinhardt and Hubbard data set has also been used to test and propose a closest neighbor method based on the composition of functional domains. Lu et al.’s Proteome Analyst is based on the annotation of homologous proteins and SWISS-PROT keywords. The LOCkey and LOChom, which were outlined by Nair and Rost in 2002, are comparable to this approach. A new technique called PSLT predicts ten subcellular localizations. One of the earliest techniques created for subcellular localization prediction was PSORT, which was introduced in 1992. PSORT is regarded as a hybrid technique since it makes use of motifs, N-terminal targeting sequence information. http://www.shanlaxjournals.com 147 £UP® Shanlax International Journal of Arts, Science and Humanities This approach predicts 14 animal and 17 plant subcellular locations using a set of knowledge-based “if-then” principles. The PSORT approach has been extended with PSORT-B and PSORT II. ESLpred is an SVM-based technique that integrates PSIBLAST scores and dipeptide composition, and it was created using the Reinhardt and Hubbard data set. A technique that takes into account details regarding sequence motifs, general sequence characteristics (such as isoelectric points and surface composition), and mRNA expression levels was introduced by Drawid and Gerstein (60). Their approach, which was tested on the yeast genome, is based on a Bayesian prediction model. A specific technique for predicting mitochondrial proteins, MITOPRED is based on amino acid content . A few of the techniques discussed are accessible as online prediction services. Table 1 displays a comprehensive set of techniques, related URLs, and references. It is impossible to compare all of the methods against one another because they differ in their localization coverage and methods for evaluating correctness. The discussion part below examines prediction accuracy concerns. Discussion Prediction accuracy can be viewed as the individual accuracy for each predicted localization or as the total accuracy for a procedure. It frequently happens that while some localizations may be predicted with a reasonable degree of precision, others cannot. Furthermore, a fair benchmark comparison is a difficult undertaking because the majority of algorithms have been trained using distinct data sets or training procedures. Although targeting sequence-based methods like TargetP and iPSORT often only predict four plant and three non-plant localizations (limited coverage), their prediction accuracy is comparatively good. It should be noted, therefore, that SP and other projected categories are not subcellular localizations. Proteins from at least six distinct subcellular localizations are included in the SP category, whereas others belong to at least three. The query protein can only be allocated a specific location if the prediction is mitochondrial or chloroplast. In certain situations, using a more specialized technique that accurately predicts mitochondrial proteins, like MITOPRED, might even be a wise decision. The difficulty of detecting the presence of a targeting sequence complicates predictions based only on targeting sequences (63). The coverage of methods based on the total makeup of amino acids varies greatly. The avalanche of various computational strategies used to the data set by Reinhardt and Hubbard, where four localizations are depicted (32–35, 51, 59), was likely sparked by the relatively basic biological model at its core. With just slight variations, practically every approach performs similarly on this data set, which has a high degree of sequence homology (up to 90%). Due to the use of various cross-validation techniques, some algorithms are more likely than others to overfit. Since very little is known about the protein in question, other techniques in this category with more localization coverage have equivalent overall accuracies and letter choice. Direct sequence homology-based methods can be highly accurate in some situations. Finding a very comparable protein with a known subcellular localization annotation is the foundation of these techniques. The events in the sorting process are also a result of predictions that ignore protein-specific properties that can be discovered from a training data set. The disadvantage is that the outcome is left up to chance if no homologous protein with indicated localization is available. When there is minimal information available about the protein of interest, hybrid approaches are the preferred approach since they typically predict a greater variety of sub cellular localizations. Since these may provide comprehensive information about potentially identified motifs and targeting sequences, methods that provide a verbose output of the prediction findings can be suggested. Protein sorting is a crucial step in simulating a tiny aspect of the cell’s systems biology, and a smart prediction approach should aim to replicate this biological process. The frequent use of machine learning in this field further demonstrates its significance in the creation and deployment of the intricate underlying biological models. 148 £UP® kupah fiy kw;Wk; mwptpay; (kfspu;) fy;Y}up> jpUney;Ntyp gz;ila tuyhWk; jkpou; gz;ghLk; Without a doubt, this sector will see a plethora of innovative techniques in the near future, both in terms of algorithms and biological motives. Sub cellular localization prediction is probably going to be a fundamental component of systems biology methodologies, which seek to comprehend the more comprehensive facets of molecular biology. References 1. Eisenberg, D., et al. 2000. Protein function in the post-genomic era. Nature 405: 823-826. 2. Koonin, E.V. 2000. Bridging the gap between sequenceand function. Trends Genet. 16: 16. 3. Shurety, W., et al. 2000. Localization and postGolgi trafficking of tumor necrosis factor-alpha in macrophages. J. Interferon Cytokine Res. 20: 427-438. 4. Bryant, D.M. and Stow, J.L. 2004. The ins and outs of E-cadherin trafficking. Trends Cell Biol. 14: 427-434. 5. Hartmann, T., et al. 1996. Alzheimer’s disease betaA4 protein release and amyloid precursor protein sorting are regulated by alternative splicing. J. Biol. Chem. 271: 13208-13214. 6. Nakai, K. 2000. Protein sorting signals and prediction of subcellular localization. Adv. Protein Chem. 54: 277-344. 7. Doerks, T., et al. 1998. Protein annotation: detective work for function prediction. Trends Genet. 14:248-250. 8. Dyall, S.D., et al. 2004. Ancient invasions: from endosymbionts to organelles. Science 304: 253257. 9. Cline, K. and Henry, R. 1996. Import and routing of nucleus-encoded chloroplast proteins. Annu. Rev.Cell. Dev. Biol. 12: 1-26. 10. Schatz, G. 1998. Protein transport. The doors to organelles. Nature 395: 439-440. 11. Palade, G. 1975. Intracellular aspects of the process of protein synthesis. Science 189: 347358. 12. Lee, M.C., et al. 2004. Bi-directional protein transport between the ER and Golgi. Annu. Rev. CellDev. Biol. 20: 87-123. 13. Neumann, U., et al. 2003. Protein transport in plant cells: in and out of the Golgi. Ann. Bot. 92: 167-180. 14. Rusch, S.L. and Kendall, D.A. 1995. Protein transport via amino-terminal targeting sequences: common themes in diverse systems. Mol. Membr. Biol. 12: 295-307. 15. Schatz, G. and Dobberstein, B. 1996. Common principles of protein translocation across membranes. Sci 16. Jarvis, P. and Robinson, C. 2004. Mechanisms of proteinimport and routing in chloroplasts. Curr. Biol.14: R1064-1077. 17. Hawlitschek, G., et al. 1988. Mitochondrial proteinimport: identification of processing peptidase and ofPEP, a processing enhancing protein. Cell 53:795-806. 18. Arretz, M., et al. 1991. Processing of mitochondrial precursor proteins. Biomed. Biochim. Acta 50: 403-412. 19. Shackleton, J.B. and Robinson, C. 1991. Transport of proteins into chloroplasts. The thylakoidal processing peptidase is a signal-type peptidase with stringent substrate requirements at the -3 and -1 positions. J. Biol. Chem. 266: 12152-12156. 20. Nigg, E.A., et al. 1991. Nuclear import-export: in search of signals and mechanisms. Cell 66: 15-22. 21. Dingwall, C. and Laskey, R.A. 1991. Nuclear targeting sequences—a consensus? Trends Biochem. Sci. 16: 478-481. 22. Scheiffele, P. and Fullekrug, J. 2000. Glycosylation and protein transport. Essays Biochem. 36: 27-35. 23. Bergeron, J.J., et al. 1994. Calnexin: a membranebound chaperone of the endoplasmic reticulum.Trends Biochem. Sci. 19: 124-128. 24. Paulson, J.C. 1989. Glycoproteins: what are the sugar chains for? Trends Biochem. Sci. 14: 272-276. 25. Silhavy, T.J., et al. 1983. Mechanisms of protein localization. Microbiol. Rev. 47: 313-344. 26. Clausmeyer, S., et al. 1993. Protein import into chloroplasts. The hydrophilic lumenal proteins exhibit unexpected import and sorting specificities in spite of structurally conserved transit peptides. J. Biol.Chem. 268: 1386913876. 27. Endo, T., et al. 1989. N-terminal half of a mitochondrial presequence peptide takes http://www.shanlaxjournals.com 149 £UP® Shanlax International Journal of Arts, Science and Humanities a helical conformation when bound to dodecylphosphocholine micelles: a proton nuclear magnetic resonance study. J. Biochem.106: 396-400. 28. Hammen, P.K., et al. 1994. Structure of the signal sequences for two mitochondrial matrix proteins that are not proteolytically processed upon import. Biochemistry 33: 8610-8617. 29. Andrade, M.A., et al. 1998. Adaptation of protein surfaces to subcellular location. J. Mol. Biol. 276:517-525. 30. Nakashima, H. and Nishikawa, K. 1994. Discrimination of intracellular and extracellular proteins using amino acid composition and residue pair frequencies. J. Mol. Biol. 238: 54-61.