Full text
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.42 https://doi.org/10.5281/zenodo.17394004 Bioinformatics: An Aid in Functional and Structural Annotation of Hypothetical Proteins Anagha Balakrishnan1, Aparna Gupta2, Dhrangadhariya Anjani2,3, John J. Georrge1,2* 1 Department of Bioinformatics, University of North Bengal, District-Darjeeling, West Bengal-734013, India 2 Department of Bioinformatics, Christ College, Rajkot, Gujarat, India 3 Faculty of Science, University of Geneva, Genève 4, Switzerland *Corresponding author: johnjgeorr[email protected] Abstract: Advancements in different techniques have resulted in a set of predicted proteins, defined as hypothetical proteins whose existence in vivo is unknown. An incomplete understanding of protein sequence/structure/function relationships causes many difficulties for prediction methods. In this genomic era, accurate functional annotations of the proteins encoded by more than half a million genomes sequenced are essential for genomic analysis. Despite several efforts, only half of the genes have been annotated in most completely sequenced genomes. In these situations, three-dimensional structural predictions combined with a suite of computational tools can suggest possible functions for the hypothetical proteins. Hence, an innovative in silico approach with bioinformatics tools to annotate proteins has been a tedious task compared to in vitro techniques. Thus, these computational tools have been used to annotate hypothetical proteins, and their predictive and accurate results prove to be an appropriate approach for functional annotation. This study provides an in-depth understanding of the various computational tools and databases available for highaccuracy annotation of proteins. Keywords: Annotation, Bioinformatics, Computational tools, Hypothetical protein, in silico, structure/function prediction 1. Introduction The GOLD database's projects and metadata coverage capabilities have grown dramatically over the last two years. The current update has 86,930 Sequencing Projects (SPs) and 63,311 Analysis Projects (APs) that have been added since the previous release in November 2022, which represents an increase of 15% and 14.6%, respectively, from the prior release (Mukherjee et al., 2024). With an overwhelming inflow of genome sequence data, the need for computational prediction of function and structure arises to decode genes. Conventional biochemical approaches have their fits and failures due to the high precision required for annotating gene function and a "price and period" drawback. Computational methods are highly desirable because only 50-60% of genes are functionally annotated in currently available completely sequenced genomes, posing one of the challenges of the postgenomic era. The present situation requires developing accurate and efficient bioinformatics tools to annotate proteins with their functions (de Crécy-lagard et al., 2022; Hong et al., 2019; Solanki et al., 2021). 2. Prediction Approaches The workflow below depicts the approach for consecutive annotation of proteins (Figure 1). Table 1 includes tools and databases alongside their URL required for protein annotation. 2.1 Sequence Similarity Search The outpouring of vast amounts of genomic data from sequencing centres inclines the bioinformatics community towards sequence similarity searches purported to identify potentially homologous sequences. Making their way to the spotlight are BLAST and FASTA, employed for biological sequence comparisons, including DNA, proteins, amino acids, and nucleotides from varied species along the taxonomic lines (Altschul et al., 1990; Lipman & Pearson, 1985). FASTA accelerates
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.43 https://doi.org/10.5281/zenodo.17394004 sequence similarity search by dividing query sequences into short stretches (words), followed by their organisation into tables depicting their location (Miller et al., 1991). The ideal word length for comparison is 1 or 2 for amino acid sequences and 5 or 6 for nucleotide sequences. Further, these words are matched with database sequences. Various versions of FASTA, like FASTX and FASTY, are available as options (Piñeiro & Pichel, 2023). Figure 1: Workflow depicting the approach for the functional annotation of proteins. Table 1: List of tools and databases for functional annotation of proteins. S. No Prediction Approach Tools & Databases URL 1 Sequence Similarity FASTA http://www.ebi.ac.uk/Tools/sss/fasta/ BLAST http://blast.ncbi.nlm.nih.gov/Blast.cgi 2 Domain Searching CDD http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi SMART http://smart.embl-heidelberg.de/ PROSITE http://prosite.expasy.org/ MyHits http://myhits.isb-sib.ch/ DoutFinder http://mendel.imp.ac.at/dout/ Interproscan https://www.ebi.ac.uk/interpro/search/sequence/ 3 Secondary Fold Annotation PSIPRED http://bioinf.cs.ucl.ac.uk/psipred/ 4 3D Homology Model Search Fugue https://fugue.mizuguchilab.org/fugue/ PDBBLAST http://blast.ncbi.nlm.nih.gov/Blast.cgi
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.44 https://doi.org/10.5281/zenodo.17394004 5 Protein Function Prediction ProtFun http://www.cbs.dtu.dk/services/ProtFun/ Pfam http://pfam.sanger.ac.uk/ 6 Protein Localisation TargetP-2.0 https://services.healthtech.dtu.dk/services/TargetP-2.0/ WoLF PSORT https://wolfpsort.hgc.jp/ TMHMM-2.0 https://services.healthtech.dtu.dk/services/TMHMM-2.0/ 7 Virulent Factor Prediction DFVF http://sysbio.unl.edu/DFVF/ FungalR http://fungalrv.igib.res.in/query.php VirulentPred 2.0 https://bioinfo.icgeb.res.in/virulent2/index.html VFDB http://www.mgc.ac.cn/VFs/main.htm VICMpred https://webs.iiitd.edu.in/raghava/vicmpred/ 8 Structure Prediction Modeller http://salilab.org/modeller/ Swiss-Model https://swissmodel.expasy.org/interactive Raptor X http://raptorx.uchicago.edu/ HHpred https://toolkit.tuebingen.mpg.de/tools/hhpred QUARK https://zhanggroup.org/QUARK/ I-TASSER http://zhanglab.ccmb.med.umich.edu/I-TASSER/ PHYRE2.2 https://www.sbg.bio.ic.ac.uk/phyre2/html/page.cgi?id=index The algorithm for BLAST works faster than FASTA, is considered equally sensitive, and is extensively used for sequence similarity (Baxevanis et al., 2020; E. S. Donkor et al., 2014). BLAST again confines its searches to word comparisons. Additionally, this one determines words that show significant similarities between the sequences. BLAST increases search stringency by limiting the search to the rarer, more weighty patterns in the protein and nucleic acid sequences. Different BLAST versions are available for various purposes, such as blastn, blastp,blastx, tblastx, tblastn, megablast, Igblast, etc. (Al-Fatlawi et al., 2023; John. J, 2016; Samal et al., 2021). BLAST is preferred over FASTA because it commits a single possible error per 10 kb of characters and allows the optimisation of several parameters according to users' requirements (E. Donkor et al., 2014). 2.2 Domain Search Another method that takes the lead in functional annotation after sequence similarity searches is the domain-searching approach, which uses software like CDD, SMART, Prosite, MyHits, DoutFinder, and InterProScan. NCBI's CDD (Conserved Domain Database) supplements protein sequence annotation data with information on the location and functional sites of conserved domain footprints (Marchler-Bauer et al., 2016). Being a provider of finely classified, major, and well-characterised protein domain families from available 3D protein structures and literature explorations, CDD remains one of the most active curators of protein domain annotation data. Models tracked by CDD represent most protein 3D structures, and the curators characterise novel families that emerge from protein structure determination efforts. Batch CD-search allows the computation and download of annotations for large sets of protein queries (Coordinators, 2015; Lu et al., 2019). SMART is utilised for the determination of modular architectures from a single sequence or entire genome, and hitherto, when applied to the complete genome sequence of yeast, has led to the discovery that about 6.7% of genes in the genome contained one or more signalling domains, which are certainly 350 additional from the formerly annotated ones (Schultz et al., 1998). Version 7 of SMART has the advantage of exporting the domain architecture analysis results to iTOL (phylogenetic tree viewer) for visualisation. 'metaSMART', a novel sub-resource facility of SMART, is committed to the exploration and analysis of domain architectures in various metagenomics data sets (Letunic & Bork, 2018; Letunic et al., 2021).
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.45 https://doi.org/10.5281/zenodo.17394004 PROSITE, the first annotated collection of motif descriptors, stores profiles and patterns in documentation designed to detect proteins and domains (Sigrist et al., 2002). Profiles and patterns in PROSITE entries are constructed based on multiple sequence comparisons of homologous sequences. Low complexity regions and distantly related sequences can thus be challenged, which remain unspotted by pairwise sequence alignments. PROSITE continues to review its motif data, which is currently accompanied by a complementary ProRule database. PROSITE contained generalised profiles for discriminating proteins and domains, which became trivial to use after the exponential increase in sequence data. ProRule's construction from PROSITE resulted in rules that can enhance its discriminating power, providing additional information on functionally and/or structurally critical amino acids. PROSITE maintains quality control through a cross-reference from the SwissProt knowledge base (Lee et al., 2007; Wang et al., 2021). The technical complexity of profiles, accompanied by the biological complexity, which cannot be reduced, poses a limitation of ProRule (Hulo et al., 2006; Sigrist et al., 2002; Sigrist et al., 2013). The latest version of PROSITE (released on 27-Nov-2024) contains 1311 patterns, 1395 profiles, and 1414 ProRules. MyHits server, an assembled open source, flexible service dedicated to annotation of protein sequences and their domain, signature analysis, recurrent updates with novel data, and an improved web interface (Pagni et al., 2004). Its flexibility compensates for its complexity. MyHits benefits include avoiding possible loopholes in MSA, providing recovered Multiple Sequence Alignment (MSA) through pfsearch and PSI-BLAST, and offering users various parameters for choosing databases, such as SWISS-PROT, RefSeq, ENSEMBL, trEST, trGEN, and trome. MAFFT is a novel method for calculating MSA based on the fast Fourier Transform in MyHits. Following this, identical and highly similar protein sequences are automatically classified. A user can realign matched sequences between two searches using ClustalW and T-Coffee options. A modified version of the JalView Applet allows users to edit MSAs and feed them back to the MyHits hub. Visual screening of large sets of sequences is possible using diverse Java-based applications like Dotlet, SEView, and Jalview. A significant bonus point is the inclusion of both public and private databases as resources. MyHits allows users to explore their results further in alternate formats, providing diverse tools and visualizing software, which can be selected at the hub. Other incorporations included the EMBOSS transeq program to translate a DNA sequence into the corresponding peptide sequence in any of the six frames (Pagni et al., 2007). DOUTfinder assists in detecting protein domains among the related protein sequences in lowcomplexity regions of sequence similarity. DOUTfinder is designed to evaluate sub-significant domain hits by providing a homology-backed procedure for post-filtering relevant sub-threshold hits. Database search services like Pfam, SMART, and CDD provide extremely reliable domain annotations when run with default threshold settings. Relaxed thresholds are preferred, along with the evaluation of obtained results in consecutive steps, to search for twilight zones in such sequences compared (Novatchkova et al., 2006). InterProScan merges various protein signature recognition methods from InterPro consortium member databases into a single resource (Blum et al., 2021). InterProScan searches now encompass families, domain repeats, post-translational modifications, active sites, binding sites, and conserved sites. The tool accommodates a newly developed JAVA GUI application called JIPS for visualisation and tracking obtained from repeated InterProScan searches (Syed & Upton, 2006). HAMAP (Highquality Automated and Manual Annotation of microbial Proteomes), integrated into InterProScan, contains weighted-matrix signatures similar to the PROSITE profiles database (Bolleman et al., 2020). To facilitate users with a cross-reference, InterProScan is linked to PRIAM, Reactome, KEGG, MetaCyc, and Unipathway (Jones et al., 2014). One of the breakthroughs in cross-reference is InterPro2GO mapping, which is cross-referenced over 66 million times in UniProtKB. InterProScan stands apart from other domain search tools and is widely used for domain search (Blum et al., 2021; Vaidya et al., 2020). The InterPro2GO mappings generate high-quality GO annotations to individual sequences based on a combination of experimental evidence and sequence analysis methods (Ulusoy & Doğan, 2024).
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.46 https://doi.org/10.5281/zenodo.17394004 2.3 Secondary Fold Annotation Arenas of modern biology remain solid in predicting the structure and function of novel proteins. In recent years, nascent algorithms emphasising profile-profile matching have provided considerably improved structure prediction. The ideal in silico tool for identifying secondary fold annotation is GenTHREADER. The position of prime tool for protein fold recognition, with attributes like fastness, reliability, and efficiency, certainly goes to GenTHREADER, which can be found on the PSIPRED server (Jones, 1999). It is a widely used server for protein fold prediction and predicting folds of individual protein sequences with distant homology to known sequences and folds. After employing a conventional sequence alignment algorithm, probable alignments are generated. These alignments are evaluated using methods based on threading techniques. Each threaded model obtained is then assessed using a neural network to develop a single measure of confidence for the proposed prediction. This neural network forms the backbone of GenTHREADER, trained to combine sequence alignment scores, length information, and pairwise alignment scores. The algorithm's speed is attributed to grid technology and low false-positive rates in results, making it ideal for automatic structure prediction of all proteins. It is available and incorporated with the Genomic Threading Database (GTD). GenTHREADER is considered the best tool for predicting the fold recognition of proteins (Jones, 1999; McGuffin et al., 2000; McGuffin & Jones, 2003; Miller et al., 1996). 2.4 3DHomology Model Search Conventional approaches for experimental determination of protein 3D structures are time-consuming and expensive, making it imperative to devise accurate computational prediction methods. Protein model search thus became prominent in predicting the 3D structure of proteins. Tools like FUGUE and PDB-BLAST are meant for the same. FUGUE associates query sequences with their distant homologs that have known structures using sequence-structure comparison methods. It utilises environment-specific substitution tables (a scoring matrix) and structure-dependent gap penalties (a gap penalty matrix), where scores for amino acid matching and insertions/deletions are evaluated depending on the local environment of each amino acid residue in a known structure. Given a query sequence or a sequence alignment as input, FUGUE scans a database of structural profiles obtained from HOMSTRAD, accesses the sequencestructure compatibility scores, and produces a list of potential homologues and alignments. The major standout of FUGUE is its environment-based approach to predicting amino acid positions. FUGUE can also be used as an accurate sequence-structure alignment program, which outperforms CLUSTALW (Shi et al., 2001; Vedithi et al., 2021; Williams et al., 2001). Using BLAST alongside the choice of PDB as a cross-reference database is an alternative way to search 3D homology structures. On query input, the tool refers to the PDB database, searching for similar structures related to the query (Pearce & Zhang, 2021). 2.5 Protein Function Prediction An increasing influx of sequence information resulted in the significant development of function prediction tools. Such a consequence was followed by a lack of experimental characterisation and hindrance of significant homologs, making existing databases less reliable. Hence, a few tools are used for function prediction methods, such as ProtFun and Pfam. The ProtFun approach is devised for the functional prediction of protein and enzyme class (EC) using an EC classification system from the amino acid sequence of an orphan protein (Jensen et al., 2002). ProtFun utilises functional attributes directly related to linear protein sequences of amino acids. Attributes include post-translational modifications, protein sorting, length, isoelectric point, composition of the polypeptide chain, etc. The server now expands the prediction method to cover biologically and pharmaceutically interesting categories in the Gene Ontology (GO) classifications system. The neural network approach predicts the function of novel protein sequences. Initially, all the sequence-derived characters are calculated and then presented to each of the five neural networks corresponding to each GO class (Jensen et al., 2003).
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.47 https://doi.org/10.5281/zenodo.17394004 The Pfam database is an extensive collection of protein families represented by multiple sequence alignments and hidden Markov models (HMM) (Finn et al., 2013). Pfam is bifurcated into two parts: PfamA, containing the collection of manually curated high-quality domains and families, and PfamB, generated from the protein domain database ProDom (Bru et al., 2005). Pfam scales itself with growth in the number of sequences deposited. Scalability is achieved by constructing a set of seed alignments that are used to train HMM and, in turn, are used to search any sequence database for homologs. Pfam also generates higher-level groupings of related families through intensive iteration (Finn, R. D. et al., 2013). The number of curators in Pfam has increased, and they annotate the protein function via the helpdesk provided by the tool (Punta et al., 2012; Yang & Ning, 2022). It is the best location for browsing the latest domain definitions (Heger & Holm, 2003). 2.6 Protein Localisation Prediction of sub-cellular localisation is one of the key functional features of proteins. Many tools, from TargetP, WoLF PSORT, and TMHMM, are widely used to predict protein sub-cellular localisation. TargetP, a tool designed for the simulated neural network concept, aims to predict the subcellular locations of novel proteins. It uses N-terminal sequence information to discriminate between proteins targeted for the mitochondrion, the chloroplast, the secretory pathway, and "other" localisations. Bulk predictions demonstrate an accuracy of 85% (plant) or 90% (non-plant) on redundancy-reduced test sets (Emanuelsson et al., 2000). The TargetP tool is employed in the identification of internal matrix targeting signals-like sequences (iMTS-Ls) in the mitochondrial precursor proteins (Boos et al., 2018). WoLF PSORT extends PSORT II and uses PSORT localisation features to predict protein localisation. WoLF PSORT converts protein amino acid sequences into numerical localisation features based on sorting signals, amino acid composition, and functional motifs. Next, it implements a weighted k-nearest neighbour classifier for classification based on these numerical features. Results can be predicted in two ways: a list of proteins with known localisation similar to the query protein and tables with detailed information about individual localisation features. To avoid overlearning, WoLF PSORT implements a wrapper method to select and use only the significant matches. Thus, it reduces the amount of information to be considered while interpreting the subcellular localisation of individual proteins. WoLF PSORT has demonstrated less sensitivity for Golgi and peroxisome in some instances, while being 70% sensitive for the nucleus, mitochondria, plasma membrane, extracellular, and chloroplast (Horton & Nakai, 1997; Horton et al., 2007; Jiang et al., 2021). TMHMM incorporates Hidden Markov Models into the existing algorithm for predictions of hydrophobicity, charge bias, helix lengths, and grammatical constraints. TMHMM performs the transmembrane prediction program with a 97-98% accuracy. Further, TMHMM qualifies in discriminating soluble and membrane proteins, giving an accuracy of 99% (Krogh et al., 2001). Recently, a deep learning-based tool called DeepTMHMM has been introduced with the potential to detect and predict the topology of transmembrane proteins with exceptional accuracy (Hallgren et al., 2022). 2.7 Virulent Factor Prediction Proteins that aid virulence in any organism, characterizing it as a pathogen, can be predicted using tools like DFVF, VirulentPred, VFDB, and VICMpred. The Database of Virulent Fungal Pathogens (DFVF) has been developed using the PYTHON programming language and contains information from literature exploration about fungal virulence factors and pathogenic genes. Literature fetching for database generation is performed using human intelligence to curate automated literature results. An in-house tool programmed in PYTHON is utilized to derive article titles and abstracts from the PubMed database, utilizing algorithms used by MedTAKMI. Organisms that form the normal flora of humans, such as Saccharomyces cerevisiae, may sometimes cause life-threatening diseases in immunocompromised patients. Hence, the database also focuses on genes and protein products from such flora. It provides basic information like Uniprot ID, gene symbol, taxonomy, ID, and links to several other databases (Cano et al., 2022; Lu et al., 2012; Ramesh et al., 2015).
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.48 https://doi.org/10.5281/zenodo.17394004 VirulentPred is a virulent factor predictor constructed on a bi-layer Cascade SVM. First-level classifiers are trained and optimised with diverse protein features like amino acid composition, dipeptide composition, higher order dipeptide composition, and Position-Specific Iterated BLAST (PSI-BLAST) generated Position Specific Scoring Matrices (PSSM). Similarity search-based modules were constructed using virulent and non-virulent proteins, utilizing tools like the BLAST database. Results (SVM and BLAST) from the first layer were cascaded to the second layer classifier for training and generating the final classifier. An accuracy of 81.8% is attributed to the bi-layer cascade SVM (Garg & Gupta, 2008). VirulentPred 2.0 is an improved version trained with virulent protein sequences from experimental results (Sharma et al., 2023). VFDB is a database for bacterial virulence factor prediction. The database interface is divided vertically into two panels: a collapsible menu panel on the left and a tabbed content panel on the right, both developed on the LINUX platform. The menu panel shows a tree-like organisation of all subclasses of virulent factors with direct links to each page for easy navigation. When the page loads, the menu automatically collapses into a clickable vertical bar, maximising the content panel's visible area. Data includes virulence factors from 16 crucial bacterial pathogens. Extensive literature studies were conducted using original research papers from PubMed to construct VFDB. Hence, a primary database was formed. Perl scripts were then employed to extract sequence and position information about these virulence factors from the GenBank. These were further classified using the COGS database. VFDB meets these demands by providing up-to-date, thought-provoking information and analytical tools (Chen et al., 2011; Chen et al., 2005; Zhang, 2008; Zhou et al., 2024). VICMpred is a web server that aids in broader functional classification, developed to predict the function of gram-negative bacterial proteins into categories such as virulence factors, information molecules, cellular processes, and metabolic molecules. VICM is based on Support Vector Machines trained on amino acid and dipeptide composition, achieving an overall accuracy of 52.39% for amino acids and 47.01% for dipeptides, respectively. A new way of functional classification using unique tetrapeptides found in the class of proteins was also devised. These tetrapeptides were used as the input feature for predicting the function of a protein, achieving an overall accuracy of 68.66%. A hybrid method assimilating amino acid, dipeptide composition, and tetrapeptide information demonstrates an accuracy of 70.75%. A five-fold cross-validation method was used to evaluate software performance (Saha & Raghava, 2006). 2.8 Three-Dimensional Structure Prediction The functional characterisation of proteins is incomplete without the availability of its 3D structure. Comparative modelling can sometimes provide a useful 3-D model for a protein sequence in the absence of an experimentally determined structure. Homology modelling predicts 3D models of proteins related to at least one known protein structure. Modeller, Swiss Model, Raptor X, HHpred, QUARK, and I-TASSER are 3D structure prediction tools for the same job. MODELLER is one of the extensively used tools in 3D protein structure prediction. The modeller predicts the 3D structure based on spatial restraints, including related protein structures found through sequence comparisons, NMR experiments, rules of secondary structure packing, sitedirected mutagenesis, fluorescence spectroscopy, etc. These restraints work on distances, angles, dihedral angles, pairs of dihedral angles, or pseudo atoms. Here, the tool's inputs are thresholds on spatial structure and the ligands to be modeled, and the output is the 3D structure that passes the threshold. MODELLER is available for download for most Unix/Linux systems, Windows, and Mac (Webb & Sali, 2016, 2021). Swiss-Model is a widely used web-based tool for homology modelling protein structures. It allows the prediction of three-dimensional structures based on the homology of known protein structures. It was created as an easy-to-use tool with a user-friendly interface to align the target sequence with appropriate Protein Data Bank (PDB) templates for constructing structural models. The program offers high reliability for structural predictions by automating crucial modelling tasks such as template identification, alignment optimisation, model creation, and quality assessment. Swiss-Model also incorporates comprehensive reports and visualisation tools to help interpret data for structural biology, drug design, and protein functional annotation (Araújo et al., 2023; Bordoli et al., 2009; Schwede et al., 2003).
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.49 https://doi.org/10.5281/zenodo.17394004 RaptorX is utilised for protein secondary structure prediction, template-based tertiary structure modelling, alignment quality assessment, and sophisticated probabilistic alignment sampling. It stands apart from other tools by the quality of alignment it delivers between the target sequence and one or more distantly related template proteins, aided by a novel nonlinear scoring function and a probabilistic consistency algorithm. Single-template threading, alignment quality assessment, multiple-template threading, and a fragment-free approach to free modelling are the four modules used by RaptorX. It takes RaptorX ~35 min to finish processing a sequence of 200 amino acids (Kallberg et al., 2014; Peng & Xu, 2011; Wu & Xu, 2021). HHPred is highly flexible, user-friendly, and provides excellent sensitivity. The first tool to implement pairwise comparison of profile hidden Markov models (HMMs) goes beyond the conventional profile-profile matching schemes, enhancing them with the statistical backbone. Querying the tool involves a single sequence or multiple sequence alignment, providing quick results in a simple, easy-to-read format, making tools like BLAST or PSI-BLAST easy to use. HHPred can produce pairwise query-template alignments, multiple alignments of the query with a set of templates selected from the search results, and 3D structural models that are calculated by the MODELLER software from these alignments. It provides cross-reference databases, like the PDB, SCOP, Pfam, SMART, COGs, and CDD (Pereira & Alva, 2021; Soding et al., 2005). A template-free protein structure prediction program called QUARK is illustrated and tested on the structure modelling of 145 non-homologous proteins. In the ninth Critical Assessment of Protein Structure Prediction, CASP experiment meeting, the QUARK server surpassed the second and third-best servers by 18% and 47% (Dhingra et al., 2020; Xu & Zhang, 2012; Yang et al., 2016). I-TASSER (iterative threading assembly refinement) is an online assembled platform for automated high-quality protein structure and function prediction based on the sequence-to-structureto-function concept. It received the first position among other structure prediction tools in the recent CASP7, CASP8, CASP9, and CASP10 experiments. On sequence input, it generates threedimensional (3D) atomic models from multiple threading alignments and iterative TASSER assembly simulations. Multiple threading alignments are generated using LOMETS. Further functional prediction of proteins occurs through structural comparisons with known proteins. The output of ITASSER includes 5 full-length atomic models, predicted secondary structures, ligand binding sites, GO terms with confidence scores, images of predicted proteins and ligand binding sites, etc. ITASSER currently privileges users with two user-defined restrictions, including many other flexibilities (Chen et al., 2024; Okella et al., 2020; Yang & Zhang, 2015). I-TASSER-MTD is an extended version of I-TASSER that models multi-domain structures of proteins from the peptide sequence (Zhou et al., 2022). PHYRE stands for Protein Homology/Analogy Recognition Engine. It is widely used in the biological community, allowing >150 submissions per day, and provides results in a simple user interface (Kelley & Jefferys, 2011). The First Phyre version was released in 2005, based on a profileprofile alignment algorithm based on each protein position-specific scoring matrix developed by Dr. Lawrence Kelley (Bennett-Lovsey et al., 2008; Kelley & Sternberg, 2009). The updated version of Phyre is Phyre 2, featuring a more advanced interface, a fully updated fold library, and an HHpred / HHsearch package for homology detection and batch processing. The fold library is updated weekly (Kelley & Sternberg, 2009). PHYRE 2.2 is the most recent version of PHYRE, which is based on template-based structure prediction of protein structures. The advanced feature allows users to submit their sequence, which will identify the most suitable AlphaFold model as the template. In addition, they have included representative structures for all the entries in the PDB (Powell et al., 2024). 3. Conclusion Significant results were achieved using computational biology tools, which proved crucial for the functional analysis of an organism's protein. Various protein features, such as virulent factors, cellular location, and structure, can be discovered with the help of computational tools. Analysing these tools has not only helped in functional proteomic studies, but their drawbacks encourage the development of better algorithms and tools. Incorporating various tools and databases will aid in better protein structural and functional annotation and bring genomic research to its maximum capabilities.
Recent Trends in Science and Technology-2024 Bioinformatics www.christcollegerajkot.edu.in, © Christ College, Rajkot, India ISBN: 9788197073274, Page No.50 https://doi.org/10.5281/zenodo.17394004 Funding None Data Availability Statement Not applicable. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Ethical approval Not applicable. References Al-Fatlawi, A., Menzel, M., & Schroeder, M. (2023). Is Protein BLAST a thing of the past? Nature Communications, 14(1), 8195. https://doi.org/10.1038/s41467-023-44082-5 Altschul, S. F., Gish, W., Miller, W., Myers, E. W., & Lipman, D. J. (1990). Basic local alignment search tool. Journal of Molecular Biology, 215(3), 403-410. https://doi.org/https://doi.org/10.1016/S00222836(05)80360-2 Araújo, R. S. A. d., Mendonça, F. J. B., Scotti, M. T., & Scotti, L. (2023). Protein modeling. Physical Sciences Reviews, 8(4), 567-582. https://doi.org/doi:10.1515/psr-2018-0161 Baxevanis, A. D., Baxevanis, A., Bader, G., & Wishart, D. (2020). Assessing pairwise sequence similarity: BLAST and FASTA. Bioinformatics. Hoboken: John Wiley & Sons, 45-78. Bennett-Lovsey, R. M., Herbert, A. D., Sternberg, M. J., & Kelley, L. A. (2008). Exploring the extremes of sequence/structure space with ensemble fold recognition in the program Phyre. Proteins, 70(3), 611-625. https://doi.org/10.1002/prot.21688 Blum, M., Chang, H.-Y., Chuguransky, S., Grego, T., Kandasaamy, S., Mitchell, A., . . . Raj, S. (2021). The InterPro protein families and domains database: 20 years on. Nucleic acids research, 49(D1), D344-D354. Bolleman, J., de Castro, E., Baratin, D., Gehant, S., Cuche, B. A., Auchincloss, A. H., . . . Pedruzzi, I. (2020). HAMAP as SPARQL rules—A portable annotation pipeline for genomes and proteomes. GigaScience, 9(2), giaa003. Boos, F., Muhlhaus, T., & Herrmann, J. M. (2018). Detection of Internal Matrix Targeting Signal-like Sequences (iMTS-Ls) in Mitochondrial Precursor Proteins Using the TargetP Prediction Tool. Bio Protoc, 8(17), e2474. https://doi.org/10.21769/BioProtoc.2474 Bordoli, L., Kiefer, F., Arnold, K., Benkert, P., Battey, J., & Schwede, T. (2009). Protein structure homology modeling using SWISS-MODEL workspace. Nature Protocols, 4(1), 1-13. https://doi.org/10.1038/nprot.2008.197 Bru, C., Courcelle, E., Carrere, S., Beausse, Y., Dalmar, S., & Kahn, D. (2005). The ProDom database of protein domain families: more emphasis on 3D. Nucleic Acids Res, 33(Database issue), D212-215. https://doi.org/10.1093/nar/gki034 Cano, R., Lenz, A. R., Galan-Vasquez, E., Ramirez-Prado, J. H., & Perez-Rueda, E. (2022). Gene Regulatory Network Inference and Gene Module Regulating Virulence in Fusarium oxysporum [Original Research]. Frontiers in Microbiology, 13. https://doi.org/10.3389/fmicb.2022.861528 Chen, L., Li, Q., Nasif, K. F. A., Xie, Y., Deng, B., Niu, S., . . . Xie, C. Y. (2024). AI-Driven Deep Learning Techniques in Protein Structure Prediction. International Journal of Molecular Sciences, 25(15), 8426. https://www.mdpi.com/1422-0067/25/15/8426 Chen, L., Xiong, Z., Sun, L., Yang, J., & Jin, Q. (2011). VFDB 2012 update: toward the genetic diversity and molecular evolution of bacterial virulence factors. Nucleic acids research, 40(D1), D641-D645. https://doi.org/10.1093/nar/gkr989 Chen, L., Yang, J., Yu, J., Yao, Z., Sun, L., Shen, Y., & Jin, Q. (2005). VFDB: a reference database for bacterial virulence factors. Nucleic Acids Res, 33(Database issue), D325-328. https://doi.org/10.1093/nar/gki008 Coordinators, N. R. (2015). Database resources of the National Center for Biotechnology Information. Nucleic Acids Res, 43(Database issue), D6-17. https://doi.org/10.1093/nar/gku1130 de Crécy-lagard, V., Amorin de Hegedus, R., Arighi, C., Babor, J., Bateman, A., Blaby, I., . . . Xu, J. (2022). A roadmap for the functional annotation of protein families: a community perspective. Database, 2022. https://doi.org/10.1093/database/baac062