scieee AI-readable full text Open interactive document viewer

Prediction of myocardial infarction complications

Conca, Alessandro

Abstract

L’infart de miocardi (IM), també conegut com a atac de cor, és una emergència mèdica mortal on el múscul del cor comença a morir perquè no està rebent prou flux sanguini. Això sol ser causat per un bloqueig a les artèries que subministren sang al cor. Si el flux sanguini no es restableix ràpidament, un atac cardíac pot causar danys cardíacs permanents i la mort. Després d’un infart de miocardi és possible que el pacient hagi d’enfrontar‐se a diverses complicacions. La majoria de les complicacions solen produir‐se durant les primeres setmanes després de tenir un atac de cor. Els infarts de miocardi causen tres grans problemes: disminució de la contractilitat, inestabilitat elèctrica i necrosi tisular. Actualment, disposant de grans conjunts de dades clíniques i mitjançant algorismes d’aprenentatge automàtic implementats a Python, és possible predir aquestes complicacions, per poder intervenir el més aviat possible per protegir la salut del pacient. L’objectiu d’aquest projecte és entrenar i provar un model capaç de predir, amb la màxima precisió possible, les possibles complicacions que poden afectar la salut del pacient. En aquest article, utilitzem un conjunt de dades clíniques del dipòsit d’aprenentatge automàtic de la UCI. Després d’una anàlisi de dades i una neteja de dades, es van entrenar diversos models per comparar els resultats finals. El conjunt de dades original inclou algunes dades que falten. Per tant, es van eliminar variables i pacients amb més del 30% de dades que falten (neteja) i es van entrenar alguns models de classificació. Les dades perdudes de les variables restants es van estimar mitjançant un mètode interactiu (Imputació). Atès que les classes de sortida estan desequilibrades, s’ha aplicat un nou mètode de remuestreig. Finalment, es comparen i s’analitzen tots els resultats.

Full text

PROJECT WORK Degree in Biomedical Engineering PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS Report Author: Conca, Alessandro Director: Mujica Delgado, Luis Eduardo Co‐Director: Ruiz Ordoñez, Magda Liliana Submission Date: June 2022 DISSERTATION RESUM L’infart de miocardi (IM), també conegut com a atac de cor, és una emergència mèdica mortal on el múscul del cor comença a morir perquè no està rebent prou flux sanguini. Això sol ser causat per un bloqueig a les artèries que subministren sang al cor. Si el flux sanguini no es restableix ràpidament, un atac cardíac pot causar danys cardíacs permanents i la mort. Després d’un infart de miocardi és possible que el pacient hagi d’enfrontar‐se a diverses complicacions. La majoria de les complicacions solen produir‐se durant les primeres setmanes després de tenir un atac de cor. Els infarts de miocardi causen tres grans problemes: disminució de la contractilitat, inestabilitat elèctrica i necrosi tisular. Actualment, disposant de grans conjunts de dades clíniques i mitjançant algorismes d’aprenentatge automàtic implementats a Python, és possible predir aquestes complicacions, per poder intervenir el més aviat possible per protegir la salut del pacient. L’objectiu d’aquest projecte és entrenar i provar un model capaç de predir, amb la màxima precisió possible, les possibles complicacions que poden afectar la salut del pacient. En aquest article, utilitzem un conjunt de dades clíniques del dipòsit d’aprenentatge automàtic de la UCI. Després d’una anàlisi de dades i una neteja de dades, es van entrenar diversos models per comparar els resultats finals. El conjunt de dades original inclou algunes dades que falten. Per tant, es van eliminar variables i pacients amb més del 30% de dades que falten (neteja) i es van entrenar alguns models de classificació. Les dades perdudes de les variables restants es van estimar mitjançant un mètode interactiu (Imputació). Atès que les classes de sortida estan desequilibrades, s’ha aplicat un nou mètode de remuestreig. Finalment, es comparen i s’analitzen tots els resultats. I PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS RESUMEN El infarto de miocardio (IM), también conocido como ataque cardíaco, es una emergencia médica mortal en la que el músculo cardíaco comienza a morir porque no recibe suficiente flujo de sangre. Esto generalmente es causado por un bloqueo en las arterias que suministran sangre al corazón. Si el flujo de sangre no se restablece rápidamente, un ataque cardíaco puede causar daño cardíaco permanente y la muerte. Después de un infarto de miocardio es posible que el paciente tenga que enfrentarse a diversas complicaciones. La mayoría de las complicaciones tienden a ocurrir dentro de las primeras semanas después de sufrir un ataque al corazón. Los infartos de miocardio causan tres problemas principales: disminución de la contractilidad, inestabilidad eléctrica y necrosis tisular. Hoy en día, al disponer de grandes conjuntos de datos clínicos y mediante algoritmos de aprendizaje automático implementados en Python, es posible predecir estas complicaciones para poder intervenir lo antes posible para proteger la salud del paciente. El objetivo de este proyecto es entrenar y probar un modelo capaz de predecir, con la mayor precisión posible, las posibles complicaciones que pueden afectar a la salud del paciente. En este documento, utilizamos un conjunto de datos clínicos del repositorio de aprendizaje automático de UCI. Después de un análisis y limpieza de datos, se entrenaron varios modelos para comparar los resultados finales. El conjunto de datos original incluye algunos datos faltantes. Por lo tanto, se eliminaron (limpieza) variables y pacientes con más del 30% de datos faltantes y se entrenaron algunos modelos de clasificación. Los datos perdidos de las variables restantes se estimaron utilizando un método interactivo (Imputación). Dado que las clases de salida están desequilibradas, se aplicó un método novedoso de remuestreo. Finalmente, todos los resultados son comparados y analizados. II DISSERTATION ABSTRACT Myocardial infarction(MI), also known as heart attack, is a deadly medical emergency where your heart muscle begins to die because it is not getting enough blood flow. This is usually caused by a blockage in the arteries that supply blood to your heart. If blood flow is not restored quickly, a heart attack can cause permanent heart damage and death. After a myocardial infarction it is possible that the patient has to face various complications. Most of the complications tend to occur within the first few weeks after having an heart attack. Myocardial infarcts cause three major problems: decreased contractility, electrical instability and tissue necrosis. Nowadays, having large clinical datasets available and through machine learning algorithms implemented in Python, it is possible to predict this complications, so as to be able to intervene as soon as possible to protect the patient’s health. The aim of this project is to train and test a model capable of predicting, with best accuracy as possible, the possible complications that may affect the patient’s health. In this paper we use a clinical dataset from UCI machine learning repository. After a data analysis and data cleaning, several models were trained to compare the final results. The original dataset includes some missing data. Therefore variables and patients with more than 30% of missing data were removed(cleaning), and some classification models were trained. The missed data of the remainded variables were estimated by using an interactive method(Imputation). Since the output classes are unbalanced, a novelty method of resampling was applied. Finally, all results are compared and analysed. III PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS ACKNOWLEDGMENTS IV DISSERTATION GLOSSARY FC: the functional class of angina pectoris in the last year CHD: coronary heart disease. HF: heart failure. ECG: electrocardiogram. AV: atrioventricular block. LBBB: left bundle branch block. RBBB: right bundle branch block. QRS: QRS complex in ECG. IU: international unit. ICU: intensive care unit. ESR: erythrocyte sedimentation rate. NSAID: non‐steroidal anti‐inflammatory drugs. V Contents RESUM I RESUMEN II ABSTRACT III 1 INTRODUCTION 10 1.1 Motivation 10 1.2 Objectives 11 2 DATA 12 2.1 Dataset description 12 2.2 Exploratoty data analysis 15 2.3 Data cleaning and imputation 16 3 MACHINE LEARNING METHODS FOR CLASSIFICATION 18 3.1 Classification models 18 3.1.1 Naive Bayes 18 3.1.2 Decision Tree 19 3.1.3 Random Forest 20 3.2 Models train and test 20 4 RESULTS 22 4.1 Initial results 22 4.1.1 Dataset imputed by using median and mode 22 4.1.2 Dataset imputed by using MICE algorithm 24 4.1.3 Analysis of the first results 25 4.2 Handling imbalanced dataset 25 4.3 SMOTE 26 4.4 Final results 27 4.4.1 Dataset imputed by using median and mode 27 4.4.2 Dataset imputed by using MICE 29 4.4.3 Analysis of the final results 30 VI DISSERTATION 5 CONCLUSIONS 31 REFERENCES 32 APPENDIX 33 5.1 Loading the dataset 33 5.2 Data analysis 33 5.3 Data Cleaning 34 5.4 Imputation 36 5.5 Training and testing 37 5.6 Plotting the results 38 5.7 SMOTE 39 VII List of Figures 1 Initial plot of ROE variable, with a potential wrong value of 140. 16 2 Plot of ROE variable, after the wrong value was deleted and imputed. 17 3 Visualization of Naive Bayes classifier[11]. 19 4 Graphical representation of Decision Trees models[13]. 19 5 Graphic representation of Random Forest model[14]. 20 6 Train‐test split.[8] 21 7 Train‐test split explanation. Image by Michael Galarnyk.[9] 21 8 (a) Confusion matrix of Random Forest model with the original imputed dataset. (b) Confusion matrix of Naive Bayes model with the original imputed dataset. (c) Confusion matrix of Decision Tree model with the original imputed dataset. 22 9 ROC curve with the original imputed dataset, before SMOTE 23 10 (a) Confusion matrix of Random Forest model with the given imputed dataset. (b) Confusion matrix of Naive Bayes model with the given imputed dataset. (c) Confusion matrix of Decision Tree model with the given imputed dataset. 24 11 ROC curve with the dataset imputed by MICE, before SMOTE 25 12 Undersampling and oversampling[18]. 26 13 (a) Final confusion matrix of Random Forest model with the dataset imputed by median and mode. (b) Final confusion matrix of Naive Bayes model with the dataset imputed by median and mode. (c) Final confusion matrix of Decision Tree model with the dataset imputed by median and mode. 27 14 Final ROC curve with the dataset imputed by median and mode, after SMOTE. 28 15 (a) Final confusion matrix of Random Forest model with the dataset imputed by MICE. (b) Final confusion matrix of Naive Bayes model with the dataset imputed by MICE. (c) Final confusion matrix of Decision Tree model with the dataset imputed by MICE. 29 16 Final ROC curve with the dataset imputed by MICE, after SMOTE. 30 VIII DISSERTATION Here you can find the table with the possible complications(outputs). Variable Name Meaning FIBR_PREDS Atrial fibrillation PREDS_TAH Supraventricular tachycardia JELUD_TAH Ventricular tachycardia FIBR_JELUD Ventricular fibrillation A_V_BLOK Third‐degree AV block OTEK_LANC Pulmonary edema RAZRIV Myocardial rupture DRESSLER Dressler syndrome ZSN Chronic heart failure REC_IM Relapse of the myocardial infarction P_IM_STEN Post‐infarction angina LET_IS Lethal outcome Table 2: Complications and outcomes of myocardial infarction description. 2.2 Exploratoty data analysis Exploratory data analysis is probably the most important part of a machine learning project. This analysis is crucial because it gives us all the important information we need to start working with the data, such as the correlation between some variables or the presence of missing values. There are two analysis that had been performed in this project: univariate and bivariate analysis. The first one refers to only one variable at time, and it is used for better understanding the real meaning of that variable and to check if there are ”unusual” values that need to be fixed or deleted. The second one is performed to find the relationship between each variable in the dataset and the target variable of interest (output) or two variables and finding the relationship between them. In this project the univariate analysis played a really important role because by analysing the dataset and all its variables it was possible, for example, to detect the variables with more than 30% of missing values. It is clear that this variables needed an imputation because it is unthinkable to work with a database that has so many missing values. The analysis was also important because it showed an unusual value in the variable ’ROE’:analysing the plot of the values distribution was easy to see that there were a wrong value, that was completely different to all the other recorded values (see figure 1). 15 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS Figure 1: Initial plot of ROE variable, with a potential wrong value of 140. Bivariate analysis was also performed but did not provide useful results given the poor correlation between the variables. 2.3 Data cleaning and imputation As said in the previous section, 7 variables were removed from the initial dataset as containing more than 30% of missing values. Next, 126 patients records were removed as containing more than 20% of missing values. After, the ”strange” value in the ’ROE’ variable had to be fixed: to solve this problem the value was deleted and replaced with a missing value (NaN). In this way it could be imputed with all the other missing values. After this first data cleaning, the dataset was imputed. The dataset had to be splitted in two different subdatasets: one for the variables containing real values and the other one for the ordinal and nominal variables. The division was made because the two subdatasets had to be imputed in two different ways. The first one (with the real values) was imputed using the median of each column (variable). The second one (with the ordinal and nominal values) was imputed using the most common value for each column. After the imputation, the two datasets were merged together, 16 DISSERTATION to obtain the final one. To compare, in figure 2 it can be seen the boxplot of the variable ’ROE’ after imputation. Figure 2: Plot of ROE variable, after the wrong value was deleted and imputed. As already said, this process of data cleaning and analysis has been done to obtain the first clean dataset, which will be compared with another given dataset already imputed. The second dataset was supply by the supervisors. It contains imputed values by using an iterative method named MICE. It is a multiple imputation approach that is used to replace missing data values in a data collection based on particular assumptions about the mechanism of data missingness (e.g., the data are missing at random, the data are missing completely at random).[16] 17 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS 3 MACHINE LEARNING METHODS FOR CLASSIFICATION The majority of machine learning algorithms can be divided into two categories: supervised and unsupervised methods. Supervised learning is a type of machine learning that makes use of labeled datasets. These datasets are used to train or ”supervise” algorithms so that they can accurately identify data or forecast outcomes. The model may test its accuracy and learn over time by using labeled inputs and outputs. Supervised learning is used for classification and regression. Unsupervised learning analyzes and clusters unlabeled data sets using machine learning methods. Without the need for human interaction, these algorithms uncover hidden patterns in data. Clustering, association, and dimensionality reduction are the three basic tasks that unsupervised learning models are utilized for[17]. Three different supervised machine learning models are used and trained in this project: Decision Tree, Naive Bayes and Random Forest. The three models are explained in more detail in the next section. 3.1 Classification models 3.1.1 Naive Bayes Bayesian classification methods are used to create Naive Bayes classifiers. Bayes’ theorem, which is an equation that describes the relationship between conditional probabilities of statistical data, is used in these. We’re interested in finding the likelihood of a label given some observable features in Bayesian classification, written as P(L|features). The Bayes theorem provides us how to describe this in terms of more easily computed quantities.[10] A model, called ’generative model’, is needed to compute the P(L|features)for each labels. It specifies the data’s putative random method of generation. The essential part of training a Bayesian classifier is specifying this generative model for each label. 18 DISSERTATION Figure 3: Visualization of Naive Bayes classifier[11]. 3.1.2 Decision Tree The decision tree algorithm belongs to the family of supervised machine learning algorithms. It can be used for both a classification problem as well as for regression problem. The goal of this algorithm is to create a model that predicts the value of a target variable, for which the decision tree uses the tree representation to solve the problem in which the leaf node corresponds to a class label and attributes are represented on the internal node of the tree[12]. The most significant benefit of this learning strategy is that it does not necessitate extensive data preparation. It’s also straightforward to comprehend and interpret. Figure 4: Graphical representation of Decision Trees models[13]. 19 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS 3.1.3 Random Forest Random forest is a supervised learning algorithm. It creates a ”forest” out of an ensemble of decision trees, which are commonly trained using the ”bagging” method. The bagging method’s basic premise is that combining different learning models improves the overall output. Random forest builds multiple decision trees and merges them together to get a more accurate and stable prediction. The hyperparameters of a random forest are quite similar to those of a decision tree or a bagging classifier. While growing the trees, the random forest adds more randomness to the model. When splitting a node, it looks for the best feature from a random subset of features rather than the most essential feature. As a result, there is a lot of variety, which leads to a better model[14]. Figure 5: Graphic representation of Random Forest model[14]. 3.2 Models train and test All the three models needed to be trained, and afterwards tested to verify the quality of the results. To do this the train‐test split technique was performed. The train‐test split is a technique for evaluating the performance of a machine learning algorithm. It can be used for any supervised learning technique and can be utilized for classification or regression tasks. Taking a dataset and separating 20 DISSERTATION it into two subgroups is the technique. The training dataset is the first subset, which is used to fit the model. The second subset is not used to train the model; instead, the dataset’s input element is given to the model, which then makes predictions and compares them to the predicted values. This method is used to fit it to existing data with known inputs and outputs, then create predictions for fresh cases in the future where we don’t have the expected output or goal values. The train‐test split method is used when the dataset is sufficently large. This because is important to have enough data to feed the model and to train it.[7] In this project 80% of the data were used to train the models and the remaining 20% was used for the testing part. The technique was successfully performed thanks to the Scikit‐Learn Python machine learning library provides an implementation of the train‐test split evaluation procedure via the train_test_split() function. The detailed explanation of the function can be found in the APPENDIX section. Figure 6: Train‐test split.[8] Figure 7: Train‐test split explanation. Image by Michael Galarnyk.[9] 21 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS 4 RESULTS 4.1 Initial results In this section will be presented the first result obtained: first will be shown the confusion matrix for each model. A confusion matrix is a summary of classifcaton problem predicton outcomes. The number of correct and bad predictons is totaled and broken down by class using count values. The confusion matrix’s key is this. The confusion matrix depicts the various ways in which your classifcaton model becomes perplexed when making predictons. It informs you not only about the faults made by your classifer, but also about the types of errors that are being made. Afterwards, a table with the evaluation parameters compared for each model is shown. Finally it is shown the ROC curve for the three models. It is possible to see the AUROC. AUROC(area under the Receiver Operatng Characteristc curve) is thus a performance metric for “discriminaton”: it tells about the model’s ability to discriminate between cases (positve examples) and non‐cases (negatve examples). In this section will be provided the results of the prediction about ’FIBR_PREDS’, the first of the possible complications in our dataset. The results of all the complications have been checked and it is sufficient to show only one output variable, because it is the same procedure for all complications. 4.1.1 Dataset imputed by using median and mode (a) Random Forest (b) Naive Bayes (c) Decision Tree Figure 8: (a) Confusion matrix of Random Forest model with the original imputed dataset. (b) Confusion matrix of Naive Bayes model with the original imputed dataset. (c) Confusion matrix of Decision Tree model with the original imputed dataset. 22 DISSERTATION Evaluation Parameters Naive Bayes Random Forest Decision Tree Precision 0.10622710622710622 0.0 0.4117647058823529 Accuracy 0.2012779552715655 0.8881789137380192 0.8690095846645367 Sensibility 0.8285714285714286 0.0 0.4 Specificity 0.1223021582733813 1.0 0.9280575539568345 Table 3: Performance metrics of the trained models using the dataset imputed by median and mode. Figure 9: ROC curve with the original imputed dataset, before SMOTE 23 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS 4.1.2 Dataset imputed by using MICE algorithm (a) Random Forest (b) Naive Bayes (c) Decision Tree Figure 10: (a) Confusion matrix of Random Forest model with the given imputed dataset. (b) Confusion matrix of Naive Bayes model with the given imputed dataset. (c) Confusion matrix of Decision Tree model with the given imputed dataset. Evaluation Parameters Naive Bayes Random Forest Decision Tree Precision 0.09507042253521127 0.0 0.1891891891891892 Accuracy 0.21212121212121213 0.9090909090909091 0.8393939393939394 Sensibility 0.9 0.0 0.23333333333333334 Specificity 0.14333333333333334 1.0 0.9 Table 4: Performance metrics of the trained models using the dataset imputed by MICE. 24 DISSERTATION 5 CONCLUSIONS The final results of this project are of high quality. Despite some not optimal initial results, the problem (the unbalanced dataset) was promptly identified. Thanks to the SMOTE technique it was possible to significantly improve the results, even beyond the best expectations, especially as regards the Random Forest model, which produced near‐perfect results, and the Decision Tree model. Despite this Naive Bayes model continues to have suboptimal parameters so it is not ideal for our predictions. Also Naive Bayes parameters have improved dramatically but not enough to be used for a delicate prediction like that of myocardial infarction patients. 31 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS REFERENCES [1] Golovenkin, S. E., Bac, J., Chervov, A., Mirkes, E. M., Orlova, Y. V., Barillot, E., Gorban, A. N., Zinovyev, A. (2020). Trajectories, bifurcations, and pseudo‐time in large clinical datasets: applications to myocardial infarction and diabetes data. GigaScience, 9(11). [2] University of Leicester. (2020, 7 julio). Myocardial infarction complications Database. Figshare. [3] Neural networks for forecasting of myocardial infarction complications. (1995). IEEE Conference Publication | IEEE Xplore. [4] Heart attack: What is it, causes, symptoms & treatment. Cleveland Clinic. [5] Myocardial infarction ‐ statpearls ‐ NCBI bookshelf. [6] Wikimedia Foundation. (2021, October 22). Myocardial infarction complications. Wikipedia. [7] Brownlee, J. (2020, 26 agosto). Train‐Test Split for Evaluating Machine Learning Algorithms. Machine Learning Mastery. [8] Volpi, G. F. (2021, 11 diciembre). 6 amateur mistakes I’ve made working with train‐test splits. Medium. [9] Galarnyk, M. (2022, 30 abril). Understanding Train Test Split (Scikit‐Learn + Python). Medium. [10] VanderPlas, J. (2016c). In Depth: Naive Bayes Classification | Python Data Science Handbook. The Python Data Science Handbook. [11] Alam, B. (2022, 17 abril). Implementing Naive Bayes Classification using Python. Hands‐On‐ Cloud. [12] Sharma, A. (2021, 1 marzo). Decision Tree Algorithm for Classification : Machine Learning 101. Analytics Vidhya. [13] Machine Learning Decision Tree Classification Algorithm ‐ Javatpoint. [14] Donges, N. (2022, 14 abril). Random Forest Algorithm: A Complete Guide. Built In. [15] Brownlee, J. (2021, 16 marzo). SMOTE for Imbalanced Classification with Python. Machine Learning Mastery. [16] Multiple Imputation by Chained Equations (MICE) Explained. (2019, 10 agosto). Cross Validated. [17] Supervised vs. Unsupervised Learning: What’s the Difference? (2021, 12 marzo). IBM. [18] Machine learning with oversampling and undersampling techniques, Research Gate. 32 DISSERTATION APPENDIX CODE EXPLANATIONS 5.1 Loading the dataset Loading the dataset require the use of Pandas library. After downloading the .csv file from the dataset website, it is possible to load the data in python. 5.2 Data analysis Here it is shown the univariate analysis of some variables. This commands are used to plot the values of a variable to better understand the variable meaning. This code has been applied to every variable during the data exploration and analysis. Ordinal variables: Numerical variables: 33 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS 5.3 Data Cleaning First of all, it has been checked if different columns has the same values for each rows. If yes it is possible to remove the column because it will be useless for our prediction. Checking if there are rows with the same values in each columns. If yes it is possible to remove the repeating rows, because it is like we have two or more records of the same patient, so it won’t change our output. Deleting the ’ROE’ wrong value and replacing it with a misisng value(NaN), to prepare the dataset for the imputation. 34 DISSERTATION Here the column with more then 30% of missing values are deleted. To perform this it is used the ’drop’ function and it is necessary to specify the column that need to be deleted. Here the rows with more then 20% of missing values are deleted. 35 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS 5.4 Imputation The dataset is splitted in two different subdatasets to permorf in the better way and with the better precision the imputation. The ’drop’ and ’filter’ function are used to perform this split. The two subdatasets are imputed here. The first one, with the real values, is imputed using the median, through the function median(). In the second subdataset the imputation is used using the most common value for each column. After the imputation, the subdatasets are merged together using the function merge() and using the ’ID’ variable as a reference point to join them in the correct order. 36 DISSERTATION 5.5 Training and testing The dataset is splitted in X and y. The X will contain the input values to use for the prediction, while y will contain the output variables. Then another variable is created: Y. This variable is used to filter just one output to predict and use in the train and test. The variables for the training and testing are created. There is a test_size=0.2 because the 20% of the data is used for the test, while the other 80% for the train. Here the three models are performed and fitted using the fit() function. 37 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS In the code below is performed the prediction of the models. In the Decision Tree is used the funciton predict(), which returns the actual class; for the other two models is used predict_proba(), that is used to infer the class probabilities. 5.6 Plotting the results Here the ROC curve are performed using the function roc_auc_score() and comparing the expected values with the predictions. Then the ROC curves are plotted. 38 DISSERTATION Creating and plotting the confusion matrix. Using the sklearn.metrics library it is possible to calculate the evaluation parameters for each model. All the values are then printed for a better visualization. 5.7 SMOTE To improve the results the SMOTE technique is used. It is needed to use the imblearn.over_sampling library importing SMOTE. After the smote is performed, smote.fit_resample() is used to resample X and y variable. After that 39 PREDICTION OF MYOCARDIAL INFARCTION COMPLICATIONS it is possible to obtain the improved results repeating all the before explained code, this time using a balanced dataset. 40