scieee AI-readable full text Open interactive document viewer

Development of In-silico Strategies To Advance the 3Rs in Toxicity Testing

Cisneiros e Faria Lourenco, Miguel

Abstract

Testing for developmental and reproduction toxicity (DART) hazard requires the use of large numbers of mammals (rats and rabbits), and leads to high costs for industry. Alternative test systems like amoeba’s, nematodes and zebrafish embryo’s are used with the aim to reduce the number of mammalian test animals. There is currently no overall systematic approach (to enable data mining or query ‘what if’ questions) to assess the species relevance, usefulness and applicability domain of the available model systems (conventional in vivo mammalian animal models and alternative nonmammalian systems) used to assess DART. Reliably linking data of specific genes with specific responses between organisms is essential to the development of reliable alternative non-mammalian DART test systems, as well as defining the applicability domains of these models.In this project, a new model organism was introduced in a tool that establishes relationships between DART phenotypes and underlying molecular metabolic pathways. A process of downloading and updating data files was automated and a script was created for extracting orthologous genes from unprocessed files.

Full text

Katholieke Universiteit Leuven Centre of Microbial and Plant Genetics MIGUEL CISNEIROS E FARIA LOURENC¸ O Development of In-silico Strategies To Advance the 3Rs in Toxicity Testing Erasmus+ Curricular Internship Report BSc in Bioinformatics ADVISOR Professor Maria Helena de Figueiredo Ramos Caria, PhD SUPERVISOR Professor Vera van Noort, PhD July 2021 Para a Inˆ es, por tudo To Inˆ es, for everything Acknowledgements I am a citizen of the world, known to all and to all a stranger. Erasmus of Rotterdam At the moment of my application I stated that, for me, participating in the Erasmus program is being part of Europe’s history. The accomplishment of this dream of mine and the success of my internship would not be possible without the support of many people and institutions, to whom I here properly acknowledge and thank. To Professor Helena Caria, I give my warmest thanks for all the support given before and during this internship, for always being in touch when I was away and for supporting my path during these 3 years of my bachelors degree. To Professor Vera van Noort, for the unique opportunity of being able to work in the CSB research group and being able to participate in the groundbreaking DARTpaths project. To Diksha Bhalla, to my colleagues at the CSB group and to the members of DARTpaths project for participating in the daily life of my Erasmus mobility. To Professor Francisco Pina Martins, for being what I call ’bioinformatics in person’, for being a mentor, for recognizing value in me and for supporting my endeavours all the way. To the past and present coordinators of the bioinformatics bachelors degree (Professors Ant´ onio Gonc¸alves, Gabriela Gomes, Helena Caria, Raquel Barreira, Vitor Barbosa and S´ onia Santos) and all the other Professors I have met during this time of my life in Barreiro. To Professor Gabriela Gomes for putting up with me when I bothered about the Erasmus mobility, scholarships and other issues, and to Professor Marta Justino for teaching me subjects in a way so captivating that I have decided to keep studying in my masters and for motivating me into i giving this last push on my report. To the head of Escola Superior de Tecnologia do Barreiro do Instituto Polit´ ecnico de Set´ ubal, Professor Pedro Neto, for doing everything you can to elevate this school, that I can call home. To the friends and colleagues that have shared these years with me and, in particular, to those who kept in touch and helped me during my stay abroad. To Diogo Almeida, my best friend. To Jo˜ ao Pedro Bastos and Liliana Cangueiro for being my helping hands in Leuven. Last but not least, I thank my brothers Inˆ es, Pedro and Rita. You were and are my counselors and guidance, my support net, my family. For the financial support, in the form of scholarships or monetary complements, I hereby proper thank the Agˆ encia Nacional Erasmus+, the Santander bank, and the Direc¸ ˜ ao-Geral do Ensino Superior. ii Contents Acknowledgements i List of Figures iv List of Acronyms and Abreviations v Resumo vi Summary vii 1 Introduction 1 1.1 Goals of the Internship . . . . . . . . . . . . . . . . . . . . . 1 1.2 Characterization of the Hosting Institution . . . . . . . . . . 1 1.3 Internship schedule and Developed tasks . . . . . . . . . . 2 2 State of the Art / Fundamental Theory 4 2.1 Biological Background . . . . . . . . . . . . . . . . . . . . . 4 2.1.1 Developmental and Reproduction Toxicity (DART) . 4 2.1.2 The 3R Principles for Ethical Testing and Alternative TestingMethods .................... 5 2.1.3 Adverse Outcome Pathways, Orthology and Model Organisms (Drosophila melanogaster)........ 7 2.2 Computational Background . . . . . . . . . . . . . . . . . . 8 2.2.1 Python.......................... 8 2.2.2 Git............................ 8 2.2.3 Bash........................... 8 2.2.4 Notung ......................... 9 3 Materials and Methods 9 3.1 Insertion of Drosophila melanogaster as a new model organism ............................. 10 iii 3.2 Automation of the download/update of data files . . . . . . . 11 3.3 Orthologous genes extraction from Notung files + Summarizationofresults........................ 12 4 Results and Discussion 13 5 Conclusions 17 6 Conclus˜ oes 19 References 21 A Appendices 24 A.1 Appendix I - List of species specific data sources used in the Phenotype Enrichment Tool . . . . . . . . . . . . . . . . 24 A.2 Appendix II - Example of the code for downloading and updating the data files for D. melanogaster .......... 25 A.3 Appendix III - Code of the Notung Ortholog Extractor . . . . 28 iv List of Figures 1 Graphical identities of CMPG (left) and KU Leuven (right) . 2 2 Internship schedule by tasks . . . . . . . . . . . . . . . . . 3 3 Scheme summarizing the 3R Principles [6] . . . . . . . . . 6 4 Inputs and Outputs of the Phenotype Enrichment Tool . . . 10 5 Inputs and Outputs of the Notung Ortholog Extractor . . . . 13 6 Example of the differences between the original version (top image) and the current version (bottom image) of the phenotype enrichment tool, now including D. melanogaster . . 14 iv List of Acronyms and Abreviations 3R - Reduce, Refine and Replace AOP - Adverse Outcome Pathways BASH - Bourne Again SHell DART - Developmental and Reproduction Toxicity DL - Duplication-Loss DTL - Duplication-Transfer-Loss ECHA - European Chemicals Agency FDA - Food and Drug Administration FTP - File Transfer Protocol GNU - GNU’s Not Unix! ICH - International Conference for Harmonization REACH - Registration, Evaluation, Authorisation and Restriction of Chemicals TSV - Tab Separated Values v • ICH 4.1.2. Preand postnatal development, including maternal function; • ICH 4.1.3.–Embryo-fetal development. With the motto ”no data, no market” in mind, regulatory agencies like the Food and Drug Administration (FDA) and the European Chemicals Agency (ECHA) require DART testing on new substances and chemicals via legislation (ie. REACH legislation in Europe) [2] [5]. 2.1.2 The 3R Principles for Ethical Testing and Alternative Testing Methods The 3R Principles were introduced by Russel and Burch (1959) to provide a framework for performing more humane animal research [6] [7]. The following scheme briefly defines the 3Rs: 5 Figure 3: Scheme summarizing the 3R Principles [6] As a consequence of the 3R principles, conventional methods for animal testing eventually became considered less and less appropriate. These methods also present efficiency disadvantages by being time-consuming and expensive [8]. Alternative methods for testing became popular due to the advantages comparing to conventional testing [9]: • are 3Rs compliant; • uses computational resources; • based on predictive studies (existing literature); • resorts to non-mammalian model organisms. 6 2.1.3 Adverse Outcome Pathways, Orthology and Model Organisms (Drosophila melanogaster) A new alternative method for DART testing is the inference of orthologous relations between human and non-human genes that are key in Adverse Outcome Pathways (AOP). The AOP concept provides insights to the sequence of molecular events initiated from the interaction with toxic compounds, until an adverse effect is observed as outcome. This approach helps focus on identifying the molecular pathways involved in toxicological events [10]. Regarding orthology, two genes of different species are orthologous when they have descended from a common ancestor. Orthologous genes tend to be similar in molecular and biological functions to each other. This gives insight to the relationships between species, in here between humans and the focused model organisms [11]. Drosophila melanogaster, known colloquially as the fruit fly, is an ovipositor fly that feeds primarily on unripe or ripe fruits [12]. Over the past five decades, Drosophila has become a predominant model used to understand how genes direct the development of an embryo from a single cell to a mature multicellular organism [13]. From a genomics perspective, the genome of Drosophila is 60% homologous to that of humans, less redundant, and about 75% of the genes responsible for human diseases have homologs in flies [14]. The logic behind this new alternative testing method is the following: 1. A non-human subject is exposed to a chemical substance that induces a DART adverse outcome; 2. The pathway that was triggered is identified, and so are the key genes of that pathway; 7 3. Using phylogenetic tools, orthologous relationships are mapped between human genes and the subject’s genes that interfered with the pathway; • If the genes are not orthologous - the identified key genes will most likely not be present in the human pathway and an adverse outcome may not be triggered in humans [11]. • If the genes are orthologous - those identified key genes are enriched with phenotypic information, from databases and preexisting literature, and conclusions can be drawn about the possible adverse outcomes that the chemical substance may induce in humans [11]. 2.2 Computational Background 2.2.1 Python Python is a widely used general-purpose and high level programming language. It was mainly developed for emphasis on code readability, and its syntax allows programmers to express concepts in fewer lines of code [15]. 2.2.2 Git Git is a free and open source distributed version control system designed to handle everything from small to very large projects with speed and efficiency [16]. 2.2.3 Bash Bash is the GNU Project’s shell—the Bourne Again SHell. This is an sh-compatible shell that incorporates useful features from the Korn shell (ksh) and the C shell (csh). It is intended to conform to the IEEE POSIX 8 P1003.2/ISO 9945.2 Shell and Tools standard. It offers functional improvements over sh for both programming and interactive use. In addition, most sh scripts can be run by Bash without modification [17]. 2.2.4 Notung Notung a gene tree-species tree reconciliation software package that supports duplication-loss (DL) and duplication-transfer-loss (DTL) event models with a parsimony-based optimization criterion. This software utilizes novel, efficient algorithmsfor reconstructing the history of gene duplications and losses, for rooting gene trees based on duplication/loss parsimony and for the rearrangement of weakly supported areas of gene trees [18] [19]. 3 Materials and Methods As the work developed in this internship results from the continuation of the work carried out within the scope of a master’s thesis of another student, my main raw material was the phenotype enrichment tool created by that student. That tool consists of a Python script that maps orthologous genes in pathways and statistically examines whether phenotypes are enriched or not by comparing the outcome to control or reference background, returning the enriched phenotypes caused by toxicological events as a result. Only Caenorhabditis elegans (a nematode), Danio rerio (zebrafish), Dictyostelium discoideum (a slime mould) and Mus musculus (house mouse) were used as model organisms in the initial version of the tool. The steps performed by the phenotype enrichment tool are summarized in the following points: 9 1. Calling of the tool (Python file) via terminal or within an automated pipeline with command line arguments (pathway name and ID, and path to data files folder); 2. Reading of the data files as dataframes in the python environment and selection of the columns of interest in each dataframe; 3. Mapping: pathway to human genes, human genes to orthologous genes, orthologous genes to phenotypes; 4. Execution the enrichment analysis; 5. Output of the results into datafiles and closing of the tool. The input data was obtained from the species’ specific data sources stated in Appendix I. Figure 4 outlines the inputs and outputs of this phenotype enrichment tool. Figure 4: Inputs and Outputs of the Phenotype Enrichment Tool 3.1 Insertion of Drosophila melanogaster as a new model organism For the task of adding Drosophila melanogaster to this tool, the FlyBase database was selected to obtain data on alleles and phenotypes for this species. As in the case of slime mold, there was no need to resort to another database to obtain ontologies of the phenotypes. 10 The main change done to the original python script of the tool was the addition of snippets of code for the the data reading and mapping steps (see points 2 and 3 above) for D. melanogaster. As the code structure was similar between species, some reuse of code was made to guarantee coherence between versions. 3.2 Automation of the download/update of data files In the original version of the phenotype enrichment tool, there was a need for the manual download of data files by the user from user-oriented data marts, FTP sites and species specific databases. This methodology is problematic because of the following reasons: • Laborious method for the user or system administrator; • Information gets quickly outdated; • Not recommended procedure for an automated pipeline. The solution for this was the automation of the process of download and update of data files using Python libraries. The changes in the code of the enrichment tool had to be tailor-made to the different data sources, since there was no universal approach that could be used due to the diversity of formats used by the databases to store their information. A description of the purpose of each library used is stated in the following points: •os - used for deleting older files; •datetime - used to compare the time of creation of the data files with the time when the instance of the tool started to being executed; •pathlib - file path management; 11 •wget, urllib3 - download of data files; •gzip - unzip (when necessary) of data files. A timeframe of 24 hours was set as a maximum period between the last update of data files and the execution of the tool. After that period, the functions for download of the data files and deletion of older ones are activated when running a new tool instance and a new 24 hours period starts. 3.3 Orthologous genes extraction from Notung files + Summarization of results Within the goal of the DARTpaths project of summarizing biological data for easy access and visualization by researchers, there was the need of adding orthology data of Drosophila melanogaster to the already existing data of other non-human species used as model organisms in this platform. The Notung software was used, by a colleague, to perform phylogenetic inference and extracting, in consequence, orthologous matches between human and non-human genes. Due to the long period of computation involved, the task of extracting human orthologs of D. melanogaster was never included in the batch of tasks for Notung to execute (unlike the case of the preexisting model organisms). As the output of the execution of Notung included data on orthologous relations between human genes and genes from all non-human species in a raw format, this constituted the basis for the task of extracting orthologous genes of D. melanogaster through a bioinformatical approach. A Python script was written for this, with the following workflow: 1. Calling of the script, reading of files with a specific name structure (see figure 5), and verification of the existence of a matrix with species in the columns and in the rows; 12 2. Dropping of empty values and filtering of species to only show the human orthologs of D. melanogaster; 3. Addition of the resulting data (orthologs) to a TSV (Tab Separated Values) file, summarizing all the results. 4. Repetition of the process for each file and closing of the script. Figure 5: Inputs and Outputs of the Notung Ortholog Extractor 4 Results and Discussion Introducing Drosophila as a new model organism into the phenotype enrichment tool enables further analysis on DART. This implied the addition of 33 lines of code and the change of 1 line that already existed. The addition of lines was mainly focused on the orthologous gene mapping step of this tool, as the other steps had an universal approach to all model organisms. Due to the data organization defined by FlyBase, the matching of genes with phenotypes had alleles as an intermediary step. In other words, there was a need of using python resources to connect genes to alleles and alleles to phenotypes, instead of just connecting directly genes to phenotypes. This had previously occurred with data of slime mould, so the solution for the data connection with Drosophila was replicated from the solution designed by the Masters student for slime mould. 13 The reuse of code from other model organisms enabled the assurance of coherence and continuity of the script structure of this tool, translating into efficiency. Figure 6 shows an example of the differences between the original version of the phenotype enrichment tool and the version that now includes D. melanogaster. Figure 6: Example of the differences between the original version (top image) and the current version (bottom image) of the phenotype enrichment tool, now including D. melanogaster 14 References [1] Festing, S. Wilkinson, R. (2007). The ethics of animal research. Talking Point on the use of animals in scientific research. EMBO reports, 8(6), 526–530. https://doi.org/10.1038/sj.embor.7400993 [2] European Chemicals Agency. (n.d.). Animal testing under REACH. European Chemicals Agency - ECHA. Retrieved July 20, 2021, from https://echa.europa.eu/animal-testing-under-reach [3] Nuffield Council on Bioethics. (2005). The ethics of research involving animals. London, England: Nuffield Council on Bioethics [4] EUPATI: Patient Engagement Through Education. (2020, March 27). Developmental and reproductive toxicity. EUPATI Toolbox Glossary. https://toolbox.eupati.eu/glossary/developmental-and-reproductivetoxicity/ [5] Faqi, A. S. (2012). A critical evaluation of developmental and reproductive toxicology in nonhuman primates. Systems Biology in Reproductive Medicine, 58(1), 23–32. https://doi.org/10.3109/19396368.2011.648821 [6] National Centre for the Replacement, Refinement and Reduction of Animals in Research. (n.d.). The 3Rs — NC3Rs. Retrieved July 2, 2021, from https://nc3rs.org.uk/the-3rs [7] W. Russel and B. R.L (1959). The principles of humane experimental technique, Medical Journal of Australia, vol. 1, no. 13, pp. 500, 1960. [8] Fr¨ ohlich, E. (2018). Comparison of conventional and advanced in vitro models in the toxicity testing of nanoparticles. Artificial Cells, Nanomedicine, and Biotechnology, 46(sup2), 1091–1107. https://doi.org/10.1080/21691401.2018.1479709 21 [9] Racz, P. I., Wildwater, M., Rooseboom, M., Kerkhof, E., Pieters, R., Yebra-Pimentel, E. S., Dirks, R. P., Spaink, H. P., Smulders, C., Whale, G. F. (2017). Application of Caenorhabditis elegans (nematode) and Danio rerio embryo (zebrafish) as model systems to screen for developmental and reproductive toxicity of Piperazine compounds. Toxicology in Vitro, 44, 11–16. https://doi.org/10.1016/j.tiv.2017.06.002 [10] Willett, C. (2014). Adverse Outcome Pathways. In Encyclopedia of Toxicology (pp. 95–99). Elsevier. https://doi.org/10.1016/b978-0-12386454-3.01244-6 [11] Sholihah, N. (2021, January). Phenotype Enrichment of Conserved Molecular Pathways to Advance 3R (Replacement, Reduction and Refinement) in Toxicity Testing. KULeuven. [12] Perveen, F. K. (2018). Introduction to Drosophila. In Drosophila melanogaster - Model for Recent Advances in Genetics and Therapeutics. InTech. https://doi.org/10.5772/67731 [13] Jennings, B. H. (2011). Drosophila – a versatile model in biology medicine. Materials Today, 14(5), 190–195. https://doi.org/10.1016/s1369-7021(11)70113-4 [14] Mirzoyan, Z., Sollazzo, M., Allocca, M., Valenza, A. M., Grifoni, D., Bellosta, P. (2019). Drosophila melanogaster: A Model Organism to Study Cancer. Frontiers in Genetics, 10. https://doi.org/10.3389/fgene.2019.00051 [15] Python Software Foundation. (2021, July 12). Welcome to Python. Python.Org. https://www.python.org/about/ [16] GIT-SCM. (n.d.). Git. GIT. Retrieved July 20, 2021, from https://gitscm.com/ 22 [17] Free Software Foundation. (n.d.). Bash - GNU Project - Free Software Foundation. GNU Project. Retrieved July 20, 2021, from https://www.gnu.org/software/bash/ [18] Notung Development Team. (n.d.). Notung 2.9. Carnegie Mellon University - School of Computer Science. Retrieved July 20, 2021, from https://www.cs.cmu.edu/ [19] Notung Development Team. (2019). Notung-2.9: A Manual. 23 A Appendices A.1 Appendix I - List of species specific data sources used in the Phenotype Enrichment Tool 24 A.2 Appendix II - Example of the code for downloading and updating the data files for D. melanogaster 1 2elif organism == " dmelanogaster ": 3### ENSEMBL DATABASES ( orthologs ) ##### same as the one used for orthology mapping part 4Fly_ENSEMBL_orthology_database = path_to_orthologs_database_folder+"/"+" Orthology_human_dmelanogaster_ensembl101_unique.txt" 5path = path_to_orthologs_database_folder 6xml_query = """ <? xml version ="1.0" encoding ="UTF -8"? > 7<! DOCTYPE Query > 8<Query virtualSchemaName = " default " formatter = " TSV" header = "1" uniqueRows = "1" count = "" datasetConfigVersion = "0.6" > 9 10 <Dataset name = " hsapiens_gene_ensembl " interface = " default" > 11 <Attribute name = " ensembl_gene_id " /> 12 <Attribute name = " ensembl_gene_id_version " /> 13 <Attribute name = " dmelanogaster_homolog_associated_gene_name" /> 14 <Attribute name = "dmelanogaster_homolog_ensembl_gene " /> 15 <Attribute name = " dmelanogaster_homolog_orthology_type" /> 16 </ Dataset > 17 </Query > """ 18 my_file = Path ( Fly_ENSEMBL_orthology_database ) 19 if my_file . exists (): 20 filetime = datetime . fromtimestamp (os. path . getctime ( Fly_ENSEMBL_orthology_database)) 21 if filetime < days_pass : 22 os . remove ( Fly_ENSEMBL_orthology_database ) 23 r = http . request (’GET ’,’http :// www . ensembl . org / 25 biomart / martservice ? query = ’+ xml_query ) 24 outfile = open(Fly_ENSEMBL_orthology_database , ’w ’) 25 outfile . write (r. data . decode ("utf -8")) 26 outfile . close () 27 else: 28 r = http . request (’GET ’,’http :// www . ensembl . org / biomart / martservice ? query = ’+ xml_query ) 29 outfile = open(Fly_ENSEMBL_orthology_database , ’w’) 30 outfile . write (r. data . decode ("utf -8")) 31 outfile . close () 32 ### Phenotype database 33 path = path_to_phenotype_database_folder 34 Fly_gene_to_allele_database = path_to_phenotype_database_folder+"/"+" Fly_fbal_to_fbgn_fb_current_release.tsv" 35 Fly_allele_to_phenotype_database = path_to_phenotype_database_folder+"/"+" Fly_allele_phenotypic_data_fb_current_release.tsv" 36 my_file = Path ( Fly_gene_to_allele_database ) 37 if my_file . exists (): 38 filetime = datetime . fromtimestamp (os. path . getctime ( Fly_gene_to_allele_database)) 39 if filetime < days_pass : 40 os . remove ( Fly_gene_to_allele_database ) 41 wget . download ( ’http :// ftp . flybase . org / releases / current / precomputed_files / alleles / fbal_to_fbgn_fb_FB2021_02.tsv.gz’, out = Fly_gene_to_allele_database) 42 else: 43 wget . download (’http :// ftp . flybase . org / releases / current / precomputed_files / alleles / fbal_to_fbgn_fb_FB2021_02 . tsv .gz ’, out=Fly_gene_to_allele_database) 44 45 my_file = Path(Fly_allele_to_phenotype_database) 46 if my_file . exists (): 47 filetime = datetime . fromtimestamp (os. path . getctime ( 26 Fly_allele_to_phenotype_database)) 48 if filetime < days_pass : 49 os.remove(Fly_allele_to_phenotype_database) 50 wget . download ( ’http :// ftp . flybase . org / releases / current / precomputed_files / alleles / allele_phenotypic_data_fb_2021_02.tsv.gz’, out = Fly_allele_to_phenotype_database) 51 else: 52 wget . download (’http :// ftp . flybase . org / releases / current / precomputed_files / alleles / allele_phenotypic_data_fb_2021_02.tsv.gz’, out = Fly_allele_to_phenotype_database) 27 A.3 Appendix III - Code of the Notung Ortholog Extractor 1 2### This tool extracts D. Melanogaster orthologs from Human genes 3### by searching each Notung file , then placing the orthologs in a dataframe and finally by appending that that to a TSV file . 4### If a TSV file already exists , it will be eliminated when running this script . 5 6#Library installation 7import pandas as pd 8import os 9from pathlib import Path 10 11 # States the path where the script is stored 12 path = os .path . abspath ( os.path . dirname (" __file__ ")) 13 print ( path ) 14 # Function for crawling the folders , loading the Notung file data to a dataframe , filter its content to only store the intended information ( Drosophila and Human ) , 15 # tranform the data to the intended form for output , and finally send the dataframe to the TSV file . This process occurs for each Notung file and the outputs are appended 16 #to the same TSV file. 17 18 my_file = Path (" Notung_DMelanogaster_orthologs .txt ") 19 if my_file . exists (): 20 os . remove (" Notung_DMelanogaster_orthologs . txt") 21 for root , dirs , files in os. walk (path): 22 for name in files : 23 # print ( files ) 24 if name . endswith (" ReadyForNotung . nw. gtpruned . rooting .0. rearrange .0. homologs .txt"): 25 filename = os. path . abspath ( os .path . join (root , name )) 28 26 print ( filename ) 27 lines = open( filename ). readlines () 28 strToFind = "DrosophilaMelanosgaster" 29 count = len(open( filename ). readlines ( )) 30 if count > 10: 31 ortho_data = pd. read_csv ( filename , delimiter = ’\t’, sep="\t", header = 12, index_col =0) 32 33 Dmelanogaster_cols = [ col for col in ortho_data . columns if ’ DrosophilaMelanogaster ’ in col] 34 ortho_data = ortho_data [ Dmelanogaster_cols ] 35 ortho_data = ortho_data [( ortho_data == ’O’)] 36 ortho_data = ortho_data . dropna (how=" all") 37 ortho_data = ortho_data . replace ("O", pd. Series ( ortho_data . columns , ortho_data . columns )) 38 ortho_data [ Dmelanogaster_cols ] = ortho_data [ Dmelanogaster_cols ]. replace ({ ’7227. ’:’’}, regex = True ) 39 ortho_data [ Dmelanogaster_cols ] = ortho_data [ Dmelanogaster_cols ]. replace ({ ’__DrosophilaMelanogaster’:’’ }, regex = True ) 40 ortho_data . columns = range ( ortho_data . shape [1]) 41 ortho_data = ortho_data . transpose () 42 43 hSapiens_cols = [col for col in ortho_data . columns if ’HomoS ’ in col] 44 ortho_data = ortho_data [ hSapiens_cols ] 45 ortho_data . columns = ortho_data . columns . str. replace(r" 9606. " ,"") 46 ortho_data . columns = ortho_data . columns . str. replace(r" __HomoSapiens ","") 47 ortho_data = ortho_data . dropna (how=" all") 48 49 ortho_data = ortho_data . transpose () 50 ortho_data 51 print ( ortho_data ) 52 ortho_data . to_csv (" 29 Notung_DMelanogaster_orthologs.txt", header = False , mode =’a ’, sep = "\t") 30