Full text
Universidade do Minho Escola de Engenharia Samuel Gustavo Correia Nogueira Alves Dynamic application to identify percentiles related to placenta parameters according to gestational age December 2021
Universidade do Minho Escola de Engenharia Samuel Gustavo Correia Nogueira Alves Dynamic application to identify percentiles related to placenta parameters according to gestational age Master dissertation in Bioinformatics Dissertation supervised by Professor Ana Cristina da Silva Braga Professor Rosete Maria Amorim Novais Nogueira Cardoso December 2021
DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição-NãoComercial-SemDerivações CC BY-NC-ND https://creativecommons.org/licenses/by-nc-nd/4.0/
ACKNOWLEDGEMENTS Firstly, I would like to thank my main tutor Prof. Dr. Ana Cristina Braga for the unconditional help and support during the thesis. I am grateful for all the dedication shown since the beginning, always present to help when I needed. Also I would like to thank my co-tutor Prof. Dr. Rosete Nogueira for valuable help and inputs to help me. To my parents and sister for their support during the different projects throughout my life. And lastly but definitely not least my girlfriend which during all my ups and downs always had a word of love and support when I needed it the most giving me motivation to keep working to achieve my goals.
STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho.
RESUMO Durante anos, o crescimento do feto e o seu desenvolvimento eram as medidas que os médicos usavam para estimar o desenvolvimento e crescimento do mesmo. No entanto em estudos recentes tem-se vindo a provar que o desenvolvimento e crescimento da placenta é de grande importância para estudar a saúde e crescimento do feto. Para a correta análise do desenvolvimento da placenta, curvas de crescimento precisam de ser criadas para comparação com diferentes percentis. Métodos de regressão linear foram usados para a criação de curvas de crescimento para a população Portuguesa (Nogueira et al., 2019). Neste trabalho, usando os mesmos dados usados por Nogueira et al. foram criadas curvas de crescimento usando um método estatístico conhecido por regressão de quantis que permite criar curvas usando um método mais robusto que o anteriormente usado. Foi também objetivo deste trabalho a criação de uma aplicação que permita ao utilizador colocar os valores de crescimento da placenta e poder comparar com as curvas de crescimento. Para isto, duas aplicações foram criadas. A primeira denominada APP1 tenta simular as ligações a bases de dados existentes em hospitais, tentando ao máximo recriar as condições presentes num hospital. A segunda aplicação, APP2, é mais simples e só precisa de documentos .csv com os valores da placenta, e permite a sua utilização sem necessidade de criação de algum tipo de ligação (https://samuelalves.shinyapps.io/APP2/). Ambos os objetivos foram conseguidos com sucesso, permitindo um avanço na área que cada vez mais se mostra de grande utilidade. PALAVRAS-CHAVE Placenta, Curvas de Crescimento, Desenvolvimento, Shiny, Regressão de Quantis
vii ABSTRACT For years, fetal growth and development were the measures that the doctors used to estimate fetal development and growth. However, recent studies have shown that the development and growth of the placenta is of great importance for studying the health and growth of the fetus. For correct analysis of placental development, growth curves need to be created for comparison with different percentiles. Linear regression methods were used to create growth curves for the Portuguese population (Nogueira et al., 2019). In this work, using the same data used by Nogueira et al., growth curves were created using a statistical method known as quantile regression, which allows the creation of curves using a more robust method than the one previously used (Nogueira et al., 2019). It was also the objective of this work to create an application that allows the user to enter placental growth values and compare them with the growth curves. For this, two applications were created. The first one called APP1 tries to simulate the connections to existing databases in hospitals, trying as much as possible to recreate the conditions present in a hospital. The second application, APP2, is simpler and only needs .csv documents with the placenta values and allows its use without the need to create any kind of connection (https://samuelalves.shinyapps.io/APP2/). Both goals were successfully achieved, allowing for a breakthrough in the area that is increasingly proving to be of great use. KEYWORDS Placenta, Growth Curves, Development, Shiny, Quantile Regression
viii ABBREVIATIONS COVID-19 Coronavirus Disease 19 GA Gestational Age GPIO General Purpose Input/Output IoT Internet of Things IP address Internet Protocol address LED Light-emitting diode OWASP Open Web Application Security Project RAM Random-access memory RNA Ribonucleic Acid SQL Structured Query Language SSH Secure Shell UI User Interface USB Universal Serial Bus
ix CONTENTS Acknowledgements ............................................................................................................................. iv Resumo.............................................................................................................................................. vi Abstract............................................................................................................................................. vii Contents ............................................................................................................................................ ix List of Figures ..................................................................................................................................... xi List of Tables ..................................................................................................................................... xii 1. Introduction .............................................................................................................................. 13 1.1 Motivation and context ....................................................................................................... 13 1.2 Objectives ......................................................................................................................... 14 1.3 Organization ...................................................................................................................... 14 2. State of the art .......................................................................................................................... 16 2.1 Models and statistical methods .......................................................................................... 16 2.1.1 Previous placenta studies .......................................................................................... 16 2.1.2 Placental growth curves ............................................................................................. 16 2.1.3 Regression methods .................................................................................................. 17 2.2 Technological considerations ............................................................................................. 18 2.2.1 Raspberry Pi .............................................................................................................. 18 2.2.2 MariaDB .................................................................................................................... 19 2.2.3 SQL injection ............................................................................................................. 19 2.2.4 Shiny applications ...................................................................................................... 21 3. Regression methods ................................................................................................................. 21 3.1 Linear regression .............................................................................................................. 22 3.1.1 Least squares ............................................................................................................ 23 3.1.2 Quantile regression .................................................................................................... 25 4. Case of study ............................................................................................................................ 29 4.1 Dataset used ..................................................................................................................... 29 4.2 Growth curves ................................................................................................................... 34
16 2. STATE OF THE ART In this chapter it will be discussed the state of the art about placenta growth curves and its importance, after that it will be discussed the regression methods and how they are used in the creation of grow charts. This chapter explores, based on the literature, a set of definitions relevant to the study in question, namely the concept associated with placental growth curves (2.1) and regression models (2.2). The concepts associated with the development of dynamic applications in Shiny are also approached (2.3). 2.1 Models and statistical methods 2.1.1 Previous placenta studies Despite the different problems associated with the study of the placenta as previously stated on chapter 1, to work alongside the increased interest in the study of this organ, there are some studies already being done. Some studies such as the work done by Nogueira et al., aim to provide new data and insights from it of the placentas’ different measurements. Other studies have been done to allow the researchers to use the available technologies to better study the placenta while still taking in considerations the main objective and interest behind studying this organ (Salafia et al., 2005). To corroborate studies like of Salafia et al., other studies have been done to show how studies of the placenta can be done using ultrasound techniques to accompany the development of the placenta throughout the pregnancy (Isakov et al., 2018). 2.1.2 Placental growth curves According to Nogueira et al., 2019, placenta has increasingly gained interest among the researchers. These authors pointed out that one of the reasons to the increased interest is the fact that although more studies need to be done, the percentile curves are useful on the evaluation of fetal follow-up and diseases on children. Their study consisted in evaluating several placenta’s metrics such as placenta weight, placental diameters, placental thickness and the placental weight ratio and birth/placental weight ratio, to create percentile curves of the Portuguese population (Nogueira et al., 2019). There are studies that show that the placenta can be a predictor of perinatal mortality and morbidity (O’Brien et al., 2020) such as pre-eclampsia or fetal growth restriction (Turco & Moffett, 2019), and several diseases in adult life such as cardiovascular problems in patients with low birthweight and
17 relatively big placentas which is indicated with an increased placental ratio which is the ratio between placental weight to birthweight (Lao & Wong, 1999). Growth charts are important by providing a comparison to a reference allowing the professionals to visually see the current values of the fetus/baby and its growth trajectory (Fenton, 2003). The updating and creation of the percentile curves of the placenta and comparing them between regions or countries (a study in Ireland showed that the tenth centile of a previously created percentile curves created with data from North Americans for the placental weight at 39 weeks was higher than the growth chart of the Irish population’s placenta (O’Brien et al., 2020)) are important to evaluate pregnancy problems and increase the mother education and healthcare (Nogueira et al., 2019). Given the scarcity of studies to assess the parameters associated with the placenta and its development with gestational age, the bibliographic review will be addressed in terms of the evolution charts of development parameters of the fetus and/or newborn. 2.1.3 Regression methods Most of the studies associated with growth charts use regression models. Regression is a statistical method used to investigate the relationships between variables. Quantile regression is a method used to estimate functional relations between variables for all portions of a probability distribution (Cade & Noon, 2003). Through the years there have been innumerous ways to create growth charts. One of the first methods fit smoothing curves on sample quantiles of the segmented age groups. The problem with these methods arise with the fact that they are not robust (robustness is considered the resilience of statistical procedures to deviations from the assumptions of hypothetical models (Koenker & Bassett, 1978)) to outliers and a large sample is needed to estimate the percentiles with good precision (Chen, 2005). In the 19th century Quetelet created anthropometric methods to construct reference growth charts. According to the models created, growth curves usually are constructed based on the assumptions that the different measurements are normally distributed. If in one hand, adult heights from relatively homogeneous populations are close to normal, on the other hand, children’s heights are usually nonnormally distributed. Other measurements such as weight are potentially more problematic. As such several methods have been proposed to solved this problem and normalize the measurements (Wei et al., 2006). Quantile regression is a nonparametric method which is becoming more frequently used thanks to the increase in computing power to solve the previously stated problem. Its main advantage is that it is
18 independent of any distribution or transformation to normality due to the fact that it estimates the distributions directly determining the quantiles and as such it is a more direct method of representation. It is also more robust against outliers. Due to the fact that standard deviation and mean are used in normal distributions, both of these won’t be suited to use with quantile regression and as such, quantile estimates instead of Zscore the user obtains a quantile (Kiserud et al., 2018). 2.2 Technological considerations 2.2.1 Raspberry Pi The computer used to create and host the database has been the Raspberry Pi Model B. Raspberry pi consists of computers with a single board (single-board computer) that has several versions, each with different specifications produced by the Raspberry Pi Foundation https://www.raspberrypi.org/about/. To operate the Raspberry Pi, it is used the same components as a normal computer such as a keyboard and mouse for the input and a monitor to be displayed the output. It also contains the usual pieces found on normal computers to operate such as graphics card, processor, program memory (RAM). The hard drive for the Raspberry Pi is an SD card and the whole computer is powered by USB (Universal Serial Bus) (Raspberry Pi as Internet of Things hardware: Performances and Constraints). Due to several specifications and features such as GPIO (General Purpose Input/Output) pins (that allows the user to control different components such as LEDs (Light-emitting diode) and relays), its own operating system (raspbian) that is based off Linux (although there are other operating systems available that are not Linux based) and offers ‘out of the box’ several open source resources which allow the user to create several projects such as IoT (Internet of Things) devices like a computer which automates a water treatment plant (Raspberry Pi as IoT hardware: Performances and Constraints; Raspberry Pi for Automation of Water Treatment Plant). Compared with other products similar to a Raspberry Pi, other products have a better performance, although they are more expensive, and as such Raspberry Pi is good choice due to the low price and the ability to connect to other computers in a given network just like an usual computer would. From the variety of options on the operating system to use, Raspbian was the chosen one due to its optimization for the Raspberry Pi’s hardware. To interact with the Raspberry Pi without needing to setup the input (keyboard and mouse) and output (screen) peripherals, on a computer running the operating system windows 10 it was used the application
19 PuTTY which enables the user to connect and control the Raspberry Pi with a SSH (secure shell) connection. 2.2.2 MariaDB After MySQL was bought by Oracle, one of MySQL’s developers created a new database management system which he called MariaDB. Several companies already changed their database management system to MariaDB such as Wikipedia (https://diff.wikimedia.org/2013/04/22/wikipedia-adopts-mariadb/), WordPress.com or Google (https://mariadb.org/about/). Originally created as an enhanced drop-in replacement for MySQL, it now has new features not available on MySQL (https://mariadb.com/kb/en/mariadb-vs-mysqlcompatibility/). This database management system is available in the Raspbian repository and as such it was the database management system chosen to be used on this work. The schematics of the interaction of the APP1. The schematics are shown on Figure 1. 2.2.3 SQL injection As said by Balasundaram & Ramaraj (Balasundaram & Ramaraj, 2012), no database is safe from attackers. Open Web Application Security Project (OWASP), one of the main sources for guidelines on web application security for many companies related to web security, as of 2021 claims that the 3rd highest incident rate security flaws involve code injection (https://owasp.org/Top10/). Tiwari (Tiwari & Tiwari, 2015), state that, SQL, Structured Query Language is the most used language to interact with relational databases both to define and structure the database and its manipulation. SQL injection consists on using SQL statement as an input to the application with malicious intentions. These statements are created based on the pre-defined logical expressions from a pre-defined query and the injection alters it to a true or false statement. The error messages returned by a database can also Patient’s data Figure 1 - Schematics of APP1’s data interaction with the server.
20 assist the attacker and that can be useful when the attacker does not know how the databased is structured and its queries (Anley, 2002; Boyd & Keromytis, 2004). Defines the query as a collection of statements that usually return a single “result set”. The following explanation of an example of a SQL injection is from (Anley, 2002). It starts by showing a normal SQL statement: select id, forename, surname from authors With this statement it will be retrieved the id, forename and surname columns from the authors table. This statement could be restricted to a given author: select id, forename, surname, from authors where forename = ‘john’ and surname = ‘smith’ It is important to note that the string literals john and smith are delimited with single quotes. Assuming that both the forename and surname fields are being gathered from an user input, it can be injected an SQL query by changing the input values such as: Forename: jo’hn Surname: smith Then the query is created as: Select id, forename, surname from authors where forename = ‘jo’hn’ and surname = ‘smith’ When running this query an error might be returned stating that it has an incorrect syntax near hn. This happens because with the query submitted, the single quote character breaks out of the singlequoted data. The database then tried to execute hn and failed. This would be an example where nothing wrong happens if the attacker wanted to delete the author table, it could change the query to:
21 Forename: jo’; drop table authors— Surname: Related to this work, the database produced with the APP1 could be a target for SQL injections, To try to mitigate some of the attacks, the input is coerced to an integer and the first element is the chosen to complete the query from the user’s input. 2.2.4 Shiny applications Shiny is a package to use in R that allows the user to create web applications relatively easily with R code. Authors have already stated its usefulness and that it may revolutionize the way models simulations are shared for example in the pharmacometrics’ field making it accessible to bigger group of researchers or clinicians, since the ability of sharing and communicate the information and results is quite challenging, and although creating these web applications requires from the researchers’ high competency, these skills are becoming more common in these fields. It is also shown by the authors different ways that these applications can be shared such as ShinyApps.io, Shiny Server and Shiny Server Pro which allows the creators to share the web application created to a wider audience since these solutions are not local but instead are hosted by a web server that allows user from around the world to access it (Wojciechowski et al., 2015). In the bioinformatics there are a lot of shiny web applications already created. For example, the ShinyCircos that generates a circos plot which is considered one of the best ways to present genomic data, Cipr that is a web application to annotate cell clusters from single cell rna sequences (Ekiz et al., 2020; Y. Yu et al., 2018). BEST is another example that is a web application used to retrieve cross-related enzyme functional parameters and information from BRENDA (Hidalgo et al., 2020). Another one is the robvis which is a web application used to visualize risk-of-bias assessments (McGuinness & Higgins, 2021). Due to the pandemic of COVID-19 (Coronavirus Disease 19) several web applications were created that are used to model and show in different ways the data gathered from COVID-19 (Mercatelli et al., 2020; Salehi et al., 2021). 3. REGRESSION METHODS
22 As most of the statistical methodology associated with growth curves involves regression methods, this chapter will be dedicated to the main concepts associated with this methodology. 3.1 Linear regression Regression is a statistical method used to investigate the relationships between variables and can give answers related to a response variable. The classical regression (commonly known as the regression function) is focused on the expectation of variable Y conditional on the values of a set of variables X, E(Y|X) (Weisberg, 2005). This function can have different levels of complexity although it restricts a specific location of the Y conditional distribution exclusively. According with Su et al., linear regression delivers the simplest model form to model the regression function as a linear combination of predictors. This regression is used as the basis for many modern modelling tools. Linear regression is also used due to its modelling when the sample size is small or the signal is relatively weak, providing the user with positive results (Su et al., 2012). Considering the data 𝐷={(𝑦𝑖,𝑥𝑖)∶𝑖=1,…,𝑛} , where 𝑦𝑖 is the ith response, measured on a continuous scale, 𝑥𝑖=(𝑥𝑖1,…,𝑥𝑖𝑝)𝑡∈ℝ𝑝 is the associated predictor vector; and n(>>p) is the sample size, the linear model is specified as 𝑦𝑖=𝛽0+𝛽1𝑥𝑖1+⋯+𝛽𝑝𝑥𝑖𝑝+𝜀𝑖 with 𝜀𝑖~ 𝑖𝑖𝑑 𝑁(0,𝜎2), (1) for 𝑖=1,…,𝑛. In matrix form, 𝒚=𝑿𝜷+ 𝜺 with 𝜺 ~𝑁(0,𝜎2 𝑰), (2) Where 𝒚= [𝑦𝑖]𝑛×1 is the n -dimensional response vector; 𝑿=(𝑥𝑖𝑗)𝑛×(𝑝+1)with 𝑥𝑖0=1 is often called the design matrix; and 𝜺= [𝜀𝑖]𝑛×1. There are four major statistical assumptions involved in the specification of model (1) or (2): 1. Linearity where 𝜇≡[𝐸(𝑦𝑖|𝑥𝑖)]𝑛×1=𝑿𝜷; 2. Independence 𝜀𝑖′𝑠 are independent of each other; 3. Homoscedasticity 𝜀𝑖′𝑠 have equal variance 𝜎2; 4. Normality 𝜀𝑖′𝑠 are normally distributed.
23 To allow model interpretation, it can be introduced a generic notation μx to denote the conditional mean response as 𝜇𝑥=𝐸(𝑌|𝑿𝟏,…𝑿𝒑)=𝛽0+𝛽1𝑿𝟏+⋯+𝛽𝑝𝑿𝒑 (3) Where 𝑋𝑗 is an 𝑛× 1 vector for 𝑗=1,…,𝑝. As 𝜕𝜇𝑥/𝜕𝑋𝑗=𝛽𝑗 the regression parameters can be interpreted in terms of change rate: 𝛽𝑗 corresponds to the amount of change in the conditional mean response 𝜇𝑥 with one unit increase in 𝑋𝑗, since all other predictors are fixed (Su et al., 2012). According to Weisberg (Weisberg, 2005), the variance equation, 𝑉𝑎𝑟(𝑌|𝑋=𝑥)=𝜎2 (4) Is assumed to be constant and a positive unknown value. Since the variance 𝜎2 >0 the value observed on the ith response 𝑦𝑖 will usually not equal its expected value 𝐸(𝑌|𝑋=𝑥𝑖). Considering that difference, it was created a quantity called a statistical error , or 𝑒𝑖, for case 𝑖 defined implicity by 𝑦𝑖= 𝐸(𝑌|𝑋=𝑥𝑖)+𝑒𝑖 or explicitly by 𝑒𝑖=𝑦𝑖−𝐸(𝑌|𝑋=𝑥), i.e, the difference between the observed and the expected value of the response variable. This error depends on unknown parameters in the mean function, thus, are not observable quantities but random variables that correspond to the vertical distance between the point 𝑦𝑖 and the mean function 𝐸(𝑌|𝑋=𝑥) . Weisberg makes two assumptions concerning the errors that he considers important: • 𝐸(𝑒𝑖|𝑥𝑖)=0, thus if drawn a scatterplot of the 𝑒𝑖 versus the 𝑥𝑖 it would be a null scatterplot with no patterns; • Errors are all independent, so, the value of the error in one case gives no information about the error’s value of another case. Another important topic to consider is the model estimation which has several methods of estimating both 𝛽 and 𝜎2 for linear models (examples include least squares, Bayesian approach, ridge regression, maximum likelihood, robust estimation, etc.). The one it will be characterized is the least squares. 3.1.1 Least squares
24 The least squares is considered one of the most popular methods for estimating 𝛽. This method minimizes the distance from the observed response to the predicted values 𝑄(𝛽)=∑(𝑦𝑖−𝛽0−∑𝛽𝑗 𝑝 𝑗=1 𝑥𝑖𝑗)2 𝑛 𝑖=1 =(𝒚−𝑿𝜷)𝑇(𝒚−𝑿𝜷) (5) And when differentiated relatively for 𝛽 : 𝜕𝑄(𝛽) 𝜕𝛽 =−2(𝑿𝑻𝒚−𝑿𝑻𝑿𝜷) (6) And when setting this equation to zero, it is gotten the normal equation 𝑿𝑻𝒚=𝑿𝑻𝑿𝜷 Assuming that 𝑿 is a complete column of rank 𝑝, the Gram matrix 𝑿𝑻𝑿 is positive definite. Thus, the least squares estimator 𝜷 as a unique solution to the normal equation given by 𝜷 =(𝑿𝑻𝑿)−1𝑿𝑻𝒚 (7) And the vector of fitted values, 𝒚 is 𝒚 =𝑿𝜷 =𝑿(𝑿𝒕𝑿)−𝟏𝑿𝑻𝒚=𝑯𝒚 (8) Where 𝑯=𝑿(𝑿𝑻𝑿)−𝟏𝑿𝑻 is called sometimes the hat matrix or the projection matrix. The Figure 2 shows a representation of the residuals. The residuals are the basis for the criterion function to allow the obtaining of the estimators that from a geometrical point of view, are the vertical distances between the fitted line and the actual y-values. On the Figure 2, each data point is represented by a circle and the line represents the candidate ordinary least square line. The vertical lines that are drawn between the line and points are the residuals. The points that are below the line are negative results and the above ones are positive residuals.
25 Figure 2 – Representation of residuals for the OLS method. (Source: Weisberg, 2005) 3.1.2 Quantile regression Quantile regression, introduced by Koenker and Bassat is a method used to estimate functional relations between variables for all portions of a probability distribution and allows the studying of the conditional distribution of 𝑌 on 𝑋 at different location, offering a global view on the interrelations of 𝑋 and 𝑌 (Cade & Noon, 2003). As stated by Buhai, considering the ordinary quantile which has a real value random variable 𝑌 has the following distribution 𝐹(𝑦)=Pr (𝑌≤𝑦) (9) Thus for any 𝜏 ∈(0,1) ,the 𝜏-th quantile of 𝑌 is defined as 𝑄(𝜏)=𝑖𝑛𝑓{𝑦:𝐹(𝑦)≥𝜏} (10)
32 According with the plots from Figure 4 and Figure 5, it is noted that there is enough information on each gestational age value for the construction of the percentile curves for the babies of both masculine and feminine gender. On the Table 2 the descriptive statistics for the quantitative metrics, gestational age (GA), Fetalweight, Placentalweight and both Diameters 1 and 2 are shown, in specifics the minimum and maximum value, median and mean and the 1st and 3rd Quantile. Table 2 – Descriptive statistics for the quantitative measures. GA Fetalweight Placentalweight Diameter1 Diameter2 Min. 12.00 5.4 6.0 1.70 1.5 1st Q 19.00 235.5 101.0 10.00 8.0 Median 26.00 766.0 195.0 13.00 11.0 Mean 26.71 1248.7 233.2 14.48 12.8 3rd Q 35.00 2220.0 345.5 17.00 14.6 Max. 41.00 4880.0 995.0 32.00 30.00 Figure 6 illustrates the age distribution of women in histogram form as well as the density function. Through this image it can be seen that this is practically symmetric. Figure 6 – Distribution of Maternal Age.
33 The graphs in the Figure 7 and Figure 8 illustrates the associations of diameters 1 and 2 with gestational age according fetus gender. As can be seen from the graphs, both the diameter 1 and 2 increase its size as the gestational age progresses for all the sexes (Female, Male and Ambiguous). Figure 7 – Association between diameter 1 and GA. Figure 8 – Association between diameter 2 and GA.
34 4.2 Growth curves Using the previously described dataset, and using the packages quantregGrowth (which uses the package quantreg), and ggplot2 the growth chart was made. To do that, the dataset was loaded into R using the read.csv() function, then using quantregGrowth’s package gcrq function it was created an object that contains the coefficients and was later ploted using the plot.gcrq function. This was made for every chart made, only changing the gcrq function to the corresponding column. The plots created are in the following images (Figure 9 to Figure 12). These plots do not differentiate between the genders since that would split the data and thus creating a less robust plot, and due to the fact that there were no observable differences between the genders that would justify the separation. The 97th percentile on both the Figure 9 and Figure 10 show a decline on the interval of 35 to 40 weeks of gestational age. Although still not sure about the real cause, there is 2 hypotheses. From a statistical point of view, there could be the need for more data to allow to create a better curve. From a biological point of view, this could show alterations on the placenta on the late stages of the pregnancy. Figure 9 - Growth curve for diameter1 vs gestational age.
35 Figure 10 - Growth curve for diameter2 vs gestational age. Figure 11 - Growth curve for thickness vs gestational age.
36 Figure 12 - Growth curve for placental weight vs gestational age.
37 5. APPLICATION DEVELOPMENT 5.1 System architecture 5.1.1 General overview Since the objective of the application created is to allow the doctors, and other professionals in the obstetrician field to be able to create a chart for each client’s placenta on a previously created growth chart, great efforts have been made to simulate as much as possible the systems found in hospital and obstetric clinics to create an application that could be used in a real-world situation. As such, two applications have been made: one that connects to a database to get the patient’s information like it would in a hospital (which will from now on called APP1), and the second application that uses the information of the patient on a .csv file (which will from now on called APP2). Both allow the user to see the corresponding values of the placenta on a previously created placenta’s growth chart. Both applications’ code is on the attachments. 5.1.2 APP1 The app1 was as previously stated, created to as much as possible mimic what would be the setup in an hospital: Connect do a database with the patient’s information which will be retrieved and used individually by the patient’s doctor to analyse and compare to the previously created placenta’s growth chart. As such, to simulate the hospital’s database a Raspberry Pi was used to host a MariaDB database that can be accessed by the application. The application was created with Rstudio’s package Shiny and shinydashboard. 5.1.3 Packages used in the APP1 The core of the application is made with Rstudio’s package Shiny. This package as previously stated, allows the creation of webapps using R programing language. Another package used was shinydashboard which allows the user to create a shiny webapp, but the layout is already predefined thus allowing the user to create a good looking webapp. The application layout is defined by three parts: header, sidebar and body as shown in Figure 13.
38 Figure 13 – Representation of the application layout. 5.1.4 APP1 As seen in Figure 14, the workflow on how the APP1 works is shown. It is created a dataframe that will later be plotted and used in the programs (both APP1 and APP2) outside this code. Firstly, the code starts by reading the csv file using the command “read.csv” and the directory is specified. It was then created columns from the file it was opened and the first line removal since that is the intercept line created by the gcrq command. It was also created a column named “GA” and for each iteration on the for loop, creates a line with “GA” and the respective number. It is finally created the dataframe in this example “qr4d1”. The same procedure is then repeated to create the remaining dataframes.
39 Figure 14 – APP1’s workflow’s fluxogram.. The connection object is created. To do this, it is used the dbConnect from the library “RMariaDB”. The argument dbname indicates the name of the database, the host has the IP address (Internet Protocol address) of the raspberry pi, username and password are the respective arguments to create the connection. The next step is the creation of the user interface. In this case it starts by defining the sidebar that has a text input which by default has the word “ID” written that once the user starts to write it disappears. This text input is used to define what is the identification of the mother the user wants to see, that being a number with six digits. The side bar also has action button that when it is pressed it creates a new user The body has a tab that contains all the plots and tables about the diameter of the placenta. The server is created by a function with both the input and the output. After that, the output for the tbl1 is generated using the library “DT” and using its function “renderDataTable”. It is then needed to get the necessary data to be displayed. To do that, we get it using a function from the already mentioned library “RMariaDB”, “dbGetQuery” that sends a query to the data base and using the id input given by the user gets the wanted data. The argument editable is “TRUE” so the table can be edited by the user. The observe function allows the creation of a dynamic function that is waiting for a change in its state and when it does, it triggers an event. In this case the function is waiting for a change in the values of the table. When it does, creates the variable “info” which contains information about the edit done by the user, and it has 3 information:
40 • The column that can be accessed by using “info$column” which is usually incremented by 1 due to column index offset of 1; • The row that can be accessed by using “info$row”; • Value edited by the user that can be accessed by using “info$value”. The value is transformed by using the “as.numeric” function to update the database. To create the plot, firstly it is needed to get the data by sending a query to the database. It is then used the library “ggplot” to plot all the growth curves lines and the points. 5.1.5 Database As previously stated, the database was created using MariaDB which since it was created by a former MySQL developer, will have a lot of similarities. The examples in the Figure 15 are of 2 patients (111111 and 222222) and a column for the diameter 1 and diameter 2. Figure 15 – Example of database created.
41 5.1.6 APP2 code Like with APP1, the files used to make the growth curves are first uploaded . The same is done to the values of the d2’s growth curves. This UI (User Interface) does not use the package “shinydashboard” to keep it simpler to the user. As such the webapp is created using only the “shiny” package built-in functions. This UI is composed by 4 tabs: 1 tab for the plot and other for the table for both diameters. It has also a sidebar composed of 2 forms to allow the user to upload the d1 and d2 .csv file containing their values. It also has a button that allows the user to save the changes made on the files. The server starts by making the input work. The files can be uploaded in a reactive function to allow the data to be dynamically changeable during the execution of the app. This is done for both diameters. Next 2 reactive objects are created to store the values changed by the user. The next function is for when the button is pressed, change the .csv file with the corresponding changes. To this, the function from the package rhandsontable. This only happens if there is any change stored in either value1 or value2.
48 6. CONCLUSION AND FUTURE WORK It is widely accepted that the measurements of the fetus are good indicators of the current and future health of it, but as more research is being done about placentas, new discoveries are being released to the scientific community stating that the placenta and its measurements are also good indicators of the fetus’ health. With this premise in mind, scientists are creating placenta’s growth curves for different populations and the work about the Portuguese population was already published by Nogueira et al. (Nogueira et al., 2019). To give a new perspective and possibly improve the work done by Nogueira et al. and allow researchers and doctors to use the technology to their aid, the premises were set for this work: Create new growth curves using a different statistical method compared to the work already published and create a web application that allows the doctors to see the progression of the patient’s placenta’s measurements and compare them with the newly created curves. Both objectives were successfully completed. The newly created placenta’s growth curves were generated using a more robust regression method consolidating the Portuguese population’s placenta growth curves allowing the researchers and doctors to start using these on their studies and patient’s appointments to better evaluate the fetus current and future health. The two applications created were successful at their objectives that is to allow the user to create and work with the measurements gotten by the doctor allowing both the doctor and the researcher to better study the progression of the measurements compared to the growth curves created. APP1 was able to recreate what would happen in a hospital by recreating a connection to a database. APP2 on the other hand was made to be easier to set up and work with when the user is not using a hospital database. Regarding future works, besides the UI that could be eventually improved, allowing the user to change the values and see the changes immediately on the plots without needing to re-upload the file or update the application. It would also be interesting to set up the applications to a real-world situation and be used by many researchers and doctors and hear their opinions about the applications and start improving based on their insights. An important aspect to consider is the security and as such both applications should go through an analysis such as a static one to be able to detect security flaws (the main one of the APP1 would be SQL injection and as such improve the sanitization of the queries deeper than what is already done). Another interesting idea would be to use machine learning algorithms to predict different outcomes of the fetus although that, besides other constraints, would need to work with a big dataset and that might not yet be available with the size needed.
49 REFERENCES Anley, C. (2002). Advanced SQL injection in SQL server applications. NGSSoftware Insight Security Research , 25. Asgharnia, M., Esmailpour, N., Poorghorban, M., & Atrkar-Roshan, Z. (2008). Placental weight and its association with maternal and neonatal characteristics. Acta Medica Iranica, 46(6), 467–472. Balasundaram, I., & Ramaraj, E. (2012). An efficient technique for detection and prevention of SQL injection attack using ASCII based string matching. Procedia Engineering , 30 , 183–190. Barker, D. J. P., Bull, A. R., Osmond, C., & Simmonds, S. J. (1990). Fetal and placental size and risk of hypertension in adult life. British Medical Journal, 301(6746), 259–262. Beyerlein, A. (2014). Quantile regression - Opportunities and challenges from a user’s perspective. American Journal of Epidemiology , 180 (3), 330–331. Boyd, S. W., & Keromytis, A. D. (2004). SQLrand: Preventing SQL injection attacks. Proceedings of the 2nd Applied Cryptography and Network Security (ACNS) Conference , 292–302. Buhai, S. (2004). Quantile regression: Overview and selected applications. Ad-Astra-The Young Romanian Scientists’ Journal , February , 1–20. Cade, B. S., & Noon, B. R. (2003). A gentle introduction to quantile regression for ecologists. Frontiers in Ecology and the Environment , 1 (8), 412–420. Chen, C. (2005). Growth charts of body mass index (BMI) with quantile regression. Proceedings of the 2005 International Conference on Algorithmic Mathematics and Computer Science, AMCS’05 , 1 , 114–120. Ekiz, H. A., Conley, C. J., Stephens, W. Z., & O’Connell, R. M. (2020). CIPR: a web-based R/shiny app and R package to annotate cell clusters in single cell RNA sequencing experiments. BMC Bioinformatics , 21 (1), 191. Fenton, T. R. (2003). A new growth chart for preterm babies: Babson and Benda’s chart updated with recent data and a new format. BMC Pediatrics , 3 , 1–10. Hajovsky, D. B., Villeneuve, E. F., Schneider, W. J., & Caemmerer, J. M. (2020). An Alternative Approach to Cognitive and Achievement Relations Research: An Introduction to Quantile Regression. Journal of Pediatric Neuropsychology , 6 (2), 83–95. Hidalgo, J. S., Oróstica, K. Y., Sanchez–Daza, A., & Olivera–Nappa, Á. (2020). BEST: a Shiny/R webbased application to easily retrieve cross-related enzyme functional parameters and information from BRENDA. Bioinformatics , 1–2. Hindmarsh, P. C., Geary, M. P. P., Rodeck, C. H., Jackson, M. R., & Kingdom, J. C. P. (2001). Effect of Early Maternal Iron Stores on Placental Weight and Structure. Obstetric and Gynecologic Survey, 56(2), 66– 67. Isakov, K. M. M., Emerson, J. W., Campbell, K. H., Galerneau, F., Anders, A. M., Lee, Y. K., Subramanyam, P., Roberts, A. E., & Kliman, H. J. (2018). Estimated Placental Volume and Gestational Age. American Journal of Perinatology, 35(8), 748–757. Kiserud, T., Benachi, A., Hecher, K., Perez, R. G., Carvalho, J., Piaggio, G., & Platt, L. D. (2018). The World Health Organization fetal growth charts: concept, findings, interpretation, and application. American Journal of Obstetrics and Gynecology , 218 (2), S619–S629. Koenker, R., & Bassett, G. (1978). Regression Quantiles. Econometrica , 46 , 33-50. Lao, T. T., & Wong, W. M. (1999). The neonatal implications of a high placental ratio in small-forgestational age infants. Placenta , 20 (8), 723–726. McGuinness, L. A., & Higgins, J. P. T. (2021). Risk-of-bias VISualization (robvis): An R package and Shiny web app for visualizing risk-of-bias assessments. Research Synthesis Methods , 12 (1), 55–61. Mercatelli, D., Holding, A. N., & Giorgi, F. M. (2020). Web tools to fight pandemics: the COVID-19 experience. Briefings in Bioinformatics , 1–11.
50 Nogueira, R., Cardoso, P. L., Azevedo, A., Gomes, M., Almeida, C., Varela, C., Braga, A. C., & Pinto, J. C. (2019). Placental Biometric Parameters: The Usefulness of Placental Weight Ratio and Birth/Placental Weight Ratio Percentile Curves for Singleton Gestations as a Function of Gestational age. Journal of Clinical and Anatomic Pathology , 4 , 1–15. O’Brien, O., Higgins, M. F., & Mooney, E. E. (2020). Placental weights from normal deliveries in Ireland. Irish Journal of Medical Science , 189 (2), 581–583. Salafia, C. M., Maas, E., Thorp, J. M., Eucker, B., Pezzullo, J. C., & Savitz, D. A. (2005). Measures of placental growth in relation to birth weight and gestational age. American Journal of Epidemiology, 162(10), 991–998 Salehi, M., Arashi, M., Bekker, A., Ferreira, J., Chen, D.-G., Esmaeili, F., & Frances, M. (2021). A Synergetic R-Shiny Portal for Modeling and Tracking of COVID-19 Data. Frontiers in Public Health , 8 , 1–10. Su, X., Yan, X., & Tsai, C. L. (2012). Linear regression. Wiley Interdisciplinary Reviews: Computational Statistics , 4 (3), 275–294. Tiwari, Y., & Tiwari, M. (2015). A Study of SQL of Injections Techniques and their Prevention Methods. International Journal of Computer Applications , 114 (17), 31–33. Turco, M. Y., & Moffett, A. (2019). Development of the human placenta. Development (Cambridge) , 146 (22), 1–14. Wei, Y., Pere, A., Koenker, R., & He, X. (2006). Quantile regression methods for reference growth charts. Statistics in Medicine , 25 (8), 1369–1382. Weisberg, S. (2005). Applied Linear Regression (Third Edition) Hoboken, New Jersey: John Wiley & Sons, Inc. Wojciechowski, J., Hopkins, A. M., & Upton, R. N. (2015). Interactive pharmacometric applications using R and the Shiny package. CPT: Pharmacometrics and Systems Pharmacology , 4 (3), 146–159. Yu, K., Lu, Z., & Stander, J. (2003). Quantile regression: Applications and current research areas. Journal of the Royal Statistical Society Series D: The Statistician , 52 (3), 331–350. Yu, Y., Ouyang, Y., & Yao, W. (2018). ShinyCircos: An R/Shiny application for interactive creation of Circos plot. Bioinformatics , 34 (7), 1229–1231.
51 ATTACHMENTS APP1 Code library(shiny) library(shinydashboard) library(RMariaDB) library(pool) library(DBI) library(dbplyr) library(dplyr) library(tidyverse) library(ggplot2) library(plotly) library(DT) qr3d1=read.csv(Path_to_file) X0.1=qr3d1$X0.1 X0.1=X0.1[-1] X0.25 = qr3d1$X0.25 X0.25=X0.25[-1] X0.5=qr3d1$X0.5 X0.5= X0.5[-1] X0.75=qr3d1$X0.75 X0.75=X0.75[-1] X0.9=qr3d1$X0.9 X0.9=X0.9[-1] GA = c() for (i in 1:42){GA[[i]] = i} qr4d1 = data.frame(GA, X0.1, X0.25, X0.5, X0.75, X0.9) qr3d2=read.csv(Path_to_file) X0.12=qr3d2$X0.1 X0.12=X0.12[-1] X0.252 = qr3d2$X0.25
52 X0.252=X0.252[-1] X0.52=qr3d2$X0.5 X0.52= X0.52[-1] X0.752=qr3d2$X0.75 X0.752=X0.752[-1] X0.92=qr3d2$X0.9 X0.92=X0.92[-1] GA = c() for (i in 1:42){GA[[i]] = i} qr4d2 = data.frame(GA, X0.12, X0.252, X0.52, X0.752, X0.92) conn <- dbConnect( RMariaDB::MariaDB(), dbname = "DataBaseName", host = "IP", username = "UserName", password = "Password") on.exit(dbDisconnect(conn), add = TRUE) ui <- dashboardPage( dashboardHeader(), dashboardSidebar(sidebarMenu( textInput("id", "ID") )), dashboardBody( tabItem(tabName = 'diameter', fluidRow(box(DT::dataTableOutput('tbl1'), width = 5, height = 500), box(plotOutput('plt1'), width = 5, height = 500), box(DT::dataTableOutput("tbl2"), width = 5, height = 500), box(plotOutput('plt2'), width = 5, height = 500)) )) )
53 server <- function(input, output) { output$tbl1 <- DT::renderDataTable(datatable({ dbb <- dbGetQuery(conn, paste0( "SELECT GA, `", as.integer(input$id)[1],"1` FROM placenta5")) }, rownames= FALSE, colnames = c('GA', 'D1'), editable = TRUE)) observe(str(input$tbl1_cell_edit)) observeEvent(input$tbl1_cell_edit, { info = input$tbl1_cell_edit i = info$row j = info$col = info$col + 1 # column index offset by 1 v = info$value info$value1d <- as.numeric(info$value) dbGetQuery(conn, paste0( "UPDATE placenta5 SET`", as.integer(input$id)[1],"1`=",info$value1d," WHERE GA =",info$row)) }) output$tbl2 <- DT::renderDataTable(datatable({ dbb <- dbGetQuery(conn, paste0( "SELECT GA, `", input$id,"2` FROM placenta5")) }, rownames= FALSE, colnames = c('GA2', 'D2'), editable = TRUE)) observe(str(input$tbl2_cell_edit)) observeEvent(input$tbl2_cell_edit, { info = input$tbl2_cell_edit i = info$row j = info$col = info$col + 1 # column index offset by 1 v = info$value info$value2d <- as.numeric(info$value)
54 dbGetQuery(conn, paste0( "UPDATE placenta5 SET`", as.integer(input$id)[1],"2`=",info$value2d," WHERE GA =",info$row)) }) output$plt1 <- renderPlot({ dbqr1 <- data.frame(dbGetQuery(conn, paste0( "SELECT GA, `", as.integer(input$id)[1],"1` FROM placenta5 WHERE `", input$id,"1` >0"))) p <- ggplot(qr4d1, aes(x=GA))+ geom_line(aes(y=X0.1), color="darkred")+ geom_line(aes(y=X0.25), color="darkblue")+ geom_line(aes(y=X0.5), color="blue")+ geom_line(aes(y=X0.75), color="red")+geom_line(aes(y=X0.9), color="orange")+ geom_point(dbqr1, mapping =aes(x=GA, y=dbqr1[[2]])) print(p) }) output$plt2 <- renderPlot({ dbqr2 <- data.frame(dbGetQuery(conn, paste0( "SELECT GA, `", as.integer(input$id)[1],"2` FROM placenta5 WHERE `", input$id,"2` >0"))) p2 <- ggplot(qr4d2, aes(x=GA))+ geom_line(aes(y=X0.12), color="darkred")+ geom_line(aes(y=X0.252), color="darkblue")+ geom_line(aes(y=X0.52), color="blue")+ geom_line(aes(y=X0.752), color="red")+geom_line(aes(y=X0.92), color="orange")+ geom_point(dbqr2, mapping =aes(x=GA, y=dbqr2[[2]])) print(p2) }) } shinyApp(ui, server)
55 APP2 Code library(shiny) library(DT) library(tidyverse) library(ggplot2) library(plotly) library(rhandsontable) qr3d1 = read.csv(Path_to_file) X0.1 = qr3d1$X0.1 X0.1 = X0.1[-1] X0.25 = qr3d1$X0.25 X0.25 = X0.25[-1] X0.5 = qr3d1$X0.5 X0.5 = X0.5[-1] X0.75 = qr3d1$X0.75 X0.75 = X0.75[-1] X0.9 = qr3d1$X0.9 X0.9 = X0.9[-1] GA = c() for (i in 1:40) { GA[[i]] = i } qr4d1 = data.frame(GA, X0.1, X0.25, X0.5, X0.75, X0.9) qr3d2 = read.csv(Path_to_file) X0.12 = qr3d2$X0.1 X0.12 = X0.12[-1] X0.252 = qr3d2$X0.25 X0.252 = X0.252[-1] X0.52 = qr3d2$X0.5 X0.52 = X0.52[-1]
56 X0.752 = qr3d2$X0.75 X0.752 = X0.752[-1] X0.92 = qr3d2$X0.9 X0.92 = X0.92[-1] GA = c() for (i in 1:40) { GA[[i]] = i } qr4d2 = data.frame(GA, X0.12, X0.252, X0.52, X0.752, X0.92) ui <- fluidPage(headerPanel(title = "APP2"), sidebarLayout( sidebarPanel ( shiny::fileInput(inputId = "file_id1", label = "Upload d1 file"), shiny::fileInput(inputId = "file_id2", label = "Upload d2 file"), actionButton("do", "Save files") ), mainPanel(tabsetPanel( type = "tab", tabPanel("D1table", rHandsontableOutput('d1tbl')), tabPanel("D1plot",plotOutput('plt1')), tabPanel("D2table", rHandsontableOutput('d2tbl')), tabPanel("D2plot",plotOutput('plt2')) )) )) server <- function(input, output, session) { d1 = reactive({ file = input$file_id1 if (is.null(file)) { return() }
57 else { read.csv(file = file$datapath) } }) d2 = reactive({ file = input$file_id2 if (is.null(file)) { return() } else { read.csv(file = file$datapath) } }) # d1 <- reactive({ # df1() # }) # d2 <- reactive({ # df2() # }) # # values1 <- reactiveValues() values2 <- reactiveValues() observeEvent(input$do, { values1$data <- hot_to_r(input$d1tbl) values2$data <- hot_to_r(input$d2tbl) if (is.null(values1$data)) { return() } else {