Full text
António João Oliveira da Silva September 2022 UMinho | 2022 An Intelligent Decision Support System for the Analytical Laboratories of a Chemistry Industry Universidade do Minho Escola de Engenharia António João Oliveira da Silva An Intelligent Decision Support System for the Analytical Laboratories of a Chemistry Industry
September 2022 Doctorate Thesis Doctoral Program in Information Systems and Technologies Work developed under the supervision of: Paulo Cortez António João Oliveira da Silva An Intelligent Decision Support System for the Analytical Laboratories of a Chemistry Industry Universidade do Minho Escola de Engenharia
COPYRIGHT AND TERMS OF USE OF THIS WORK BY A THIRD PARTY This is academic work that can be used by third parties as long as internationally accepted rules and good practices regarding copyright and related rights are respected. Accordingly, this work may be used under the license provided below. If the user needs permission to make use of the work under conditions not provided for in the indicated licensing, they should contact the author through the RepositoriUM of Universidade do Minho. License granted to the users of this work Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International CC BY-NC-SA 4.0 https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en ii
Acknowledgements This PhD journey was an incredible challenge, which led me to meet fantastic people who taught me a lot and for whom I would like to dedicate this work. First, I would like to thank my supervisor, Professor Paulo Cortez, for all his support. Thank you for your perseverance, trust, and encouragement so that I always wanted to learn more and be more demanding with myself. Thank you for the motivation to overcome all the challenges that were encountered in this PhD. To the fascinating colleagues with whom I shared knowledge and had incredible life experiences during the years of this PhD, both in classes and group projects, as well as in the laboratories in which I worked in the ALGORITMI Center of School of Engineering of Minho University. To my co-workers at Computer Center Graphics (CCG), namely at EPMQ IT: Engineering Process Maturity and Quality domain, without ever forgetting ”my” Machine Learning team, which I witnessed at its conception, thank you for all the patience (which was a lot), and support throughout the years that I worked and did my PhD. To my co-workers at DTx - Digital Transformation CoLab, more particularly to the Software and Information Systems team, that were also always available to me. To the incredible AI4Medimaging team, who make me feel at home, and for the great patience with me mainly in this final stage of the PhD. The cooperation with organization where this PhD thesis took place was also very important. The organization’s specialists provided us the data and their business model knowledge. Finally, your feedback was valuable for the conclusion of this PhD. Thank you for all your availability. To my friends, all of them, from childhood friends, friends from high school, university, Afonsina, thank you for the comradeship and for helping me to keep focused on my goals. Thank you for all advice and support, which were essential to keep me motivated in the most difficult moments. Finally, I would like to thank to my family, who have always supported and encouraged me during this PhD project, providing everything they could to my success. iii
STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the Universidade do Minho. , (Place) (Date) (António João Oliveira da Silva) iv
Resumo Um Sistema Inteligente de Apoio à Decisão para os Laboratórios Analíticos de uma Indústria Química A Indústria 4.0 representa a quarta revolução industrial e envolve uma implementação que utiliza várias tecnologias de informação para dar suporte à produção, bem como uma monitorização em tempo real dos processos industriais. O tópico de Business Analytics é particularmente valioso neste contexto, uma vez que resulta de uma combinação de Business Intelligence com Optimização e Previsão. O objectivo é obter conhecimentos orientados por dados que podem ser úteis para ajudar na tomada de decisões sobre processos de produção. Por exemplo, Business Analytics pode ser utilizada para analisar dados históricos para ajudar a detectar e prever problemas ou falhas na produção. Outra possibilidade interessante é a previsão de ordens de procura, que pode ajudar no processo de gestão de stocks . Este trabalho de doutoramento é realizado no âmbito de um projecto de Investigação e Desenvolvimento (I&D). O principal objectivo é a investigação e implementação de um Sistema Inteligente de Apoio à Decisão (IDSS em Inglês) que utiliza técnicas de Business Analytics (Descritiva, Prescritiva e Preditiva), integrado no conceito de Indústria 4.0 e aplicado a Laboratórios Analíticos de Empresas Químicas. Inicialmente, as necessidades das empresas analisadas foram elicitadas, e posteriormente foram desenvolvidos vários módulos do IDSS com o objectivo de resolver os objectivos das empresas Químicas. O primeiro módulo estudado foi a previsão da chegada de amostras aos Laboratórios Analíticos, utilizando uma ferramenta de Auto Machine Learning (AutoML). Em seguida, foi desenvolvido um módulo para prever o consumo de materiais nos laboratórios. Este módulo incluiu três abordagens de previsão diferentes que foram comparadas, uma com um AutoML, outra utilizando a metodologia ARIMA e a última baseada num algoritmo de aprendizagem profunda ( Long Short-Term Memory em Inglês). Os melhores resultados de previsão foram obtidos através da abordagem AutoML. Finalmente, foi desenvolvido um módulo com métodos prescritivos para atribuir os instrumentos às análises a realizar, bem como o desenvolvimento de Dashboards de fácil utilização para o IDSS concebido. O sistema IDSS completo foi avaliado através de questionários e entrevistas abertas com os gestores do Laboratório Analítico. Globalmente, foi obtido um feedback positivo. Palavras-chave: Business Analytics, Chemical Laboratories, Industry 4.0, Machine Learning, Optimization, Prediction. v
Abstract An Intelligent Decision Support System for the Analytical Laboratories of a Chemistry Industry The Industry 4.0 represents the fourth industrial revolution and involves an implementation using several Information Technologies to support production, as well as a real-time monitoring of industrial processes. The topic of Business Analytics is particularly valuable in this context, since it results from a combination of Business Intelligence with Optimization and Forecasting. The objective is to obtain datadriven knowledge that can be useful to help decision making on production processes. For example, Business Analytics can be used to analyze historical data to help detect and predict problems or failures in production. Another interesting possibility is the prediction of demand orders, which can help in the process of stock management. This PhD work is carried out within the scope of a Research & Development (R&D) project. The main objective is the research and implementation of an Intelligent Decision Support System (IDSS) that uses Business Analytics techniques (Descriptive, Prescriptive and Predictive), integrated within the Industry 4.0 concept and applied to Analytical Laboratories of Chemical companies. Initially, the analyzed company needs were elicitated, and subsequently several IDSS modules were developed aiming to solve the Chemical company goals. The first studied module was the prediction of arrival of samples at the Analytical Laboratories by using an Auto Machine Learning (AutoML) tool. Next, a module was developed for predicting the consumption of materials in the laboratories. This module included three different forecasting approaches that were compared, one with an AutoML, another using the ARIMA methodology and the last based on a deep learning algorithm (Long Short-Term Memory). The best forecasting results were achieved by the AutoML approach. Finally, a module was developed with prescriptive methods to allocate the instruments to the analyses to be performed as well as the development of the friendly user Dashboards for the designed IDSS. The full IDSS system was evaluated by using questionnaires and open interviews with the Analytical Laboratory managers. Overall, a positive feedback was obtained. Keywords: Business Analytics, Chemical Laboratories, Industry 4.0, Machine Learning, Optimization, Prediction. vi
Contents List of Figures ix List of Tables x Acronyms xi 1 Introduction 1 1.1 Motivation .................................... 1 1.2 ProblemFormulation ............................... 3 1.3 Objectives .................................... 5 1.4 Research Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 1.5 Contributions................................... 8 1.6 ThesisOrganization................................ 10 2 Background 11 2.1 Business Analytics in Industry 4.0: A Systematic Literature Review . . . . . . . . . 11 2.1.1 Introduction ............................... 11 2.1.2 RelatedWork............................... 16 2.1.3 Literature Review Method . . . . . . . . . . . . . . . . . . . . . . . . . 18 2.1.4 Literature Review Analysis . . . . . . . . . . . . . . . . . . . . . . . . 20 2.1.5 Discussion................................ 38 2.1.6 Conclusions and research implications . . . . . . . . . . . . . . . . . . 39 2.2 Other Relevant Concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 2.2.1 Machine Learning (ML) . . . . . . . . . . . . . . . . . . . . . . . . . . 42 2.2.2 Auto Machine Learning (AutoML) . . . . . . . . . . . . . . . . . . . . . 42 2.2.3 Intelligent Decision Support Systems (IDSS) . . . . . . . . . . . . . . . . 44 2.2.4 Cross-Industry Standard Process for Data Mining (CRISP-DM) . . . . . . . 46 2.3 Business Analytics applied to the Chemical Industry . . . . . . . . . . . . . . . . 47 3 Methods, Experiments and Results 50 vii
Chapter 1 Introduction This chapter contains an motivation and contextualization of the problem that leads to this doctoral thesis. Then, the research objectives and methodologies are presented. Finally, the scientific contributions of this thesis are presented, and a description about the structure of this thesis is given. 1.1 Motivation Business Analytics plays an important role in several businesses. It focuses in the analysis of historical raw data in order to achieve useful and focused insights and a better understanding of the business performance areas (Krishnamoorthi & Mathew, 2018). Business Analytics is the result of combining Business Intelligence techniques with Optimization, Forecasting, Predictive Modeling and Statistical Analysis (Arnott & Pervan, 2014). Business Analytics systems are being applied in the Industry sector, and this, in conjunction with the Industry 4.0 phenomenon, is causing significant changes in this sector. Nowadays, most of the Industry is facing times of change. This change is being enabled by new techniques and technologies, including sensors and communication devices that generate Big Data and also analytic systems capable of analyzing such data, allowing to produce new insights and knowledge about the productive system. The term Industry 4.0 is used to identify this process. The German Federal Ministry of Education and Research defines the Industry 4.0 concept as: ”the flexibility that exists in valuecreating networks is increased by the application of cyber physical production systems. This enables machines and plants to adapt their behavior to changing orders and operating conditions through selfoptimization and reconfiguration... The main focus is on the ability of the systems to perceive information, to derive findings from it and to change their behavior accordingly, and to store knowledge gained from experience. Intelligent production systems and processes as well as suitable engineering methods and tools will be a key factor to successfully implement distributed and interconnected production facilities in future Smart Factories”(Shrouf et al., 2014). This PhD was developed within a three-year Research & Development (R&D) project that was funded by a private company and whose main objective relies in the creation of an integrated intelligent system 1
CHAPTER 1. INTRODUCTION Figure 1: Relations existing between the three R&D project WPs. based on state-of-the-art technology, under the Industry 4.0 concept, and that can improve the processes and the efficiency of the organization. The organization in question is from the Chemical Industry sector. A key aspect of this R&D project, and that is addressed in this PhD work, is the adaptation of Business Analytics techniques, under the Industry 4.0 context, to Chemical Analytical Laboratories. It should be noted that the full R&D project contains three main Work Package (WP), as shown on Figure 1. The “Planning and Scheduling System” aims to provide an automatic tool for scheduling laboratory tests. The “Automation System” WP2 is more linked with hardware (e.g., robotic arm automation). The goal is to automate some manual laboratory processes. Also, it will generate real-time laboratory data from sensors. The “Information System” (WP3) aims to design and implement the Laboratories Information Systems (IS), being based on a Big Data Warehouse, to collect, storage and process the data. These three main systems are expected to heavily communicate and interact. In particular, the “Automation System” (WP2) will interact with the “Information System” (WP3) by sending the sensor data information. Also, the IS (WP3) will provide the Commands and Priorities for the laboratory machines. The planning and scheduling module (WP1) will provide information about the plans and schedules to the WP3, and the latter will send the data, insights and predictions to the WP1. This thesis aims to cover the ”Intelligence”component of the IS (WP3). It will be focused on the design of a Business Analytics system that is capable of analyzing historical laboratory data, under the Industry 4.0 concept, in order to extract useful knowledge (e.g., predictions, insights) to improve laboratory 2
1.2. PROBLEM FORMULATION processes and management. 1.2 Problem Formulation Generally, an industrial context may be established in any one of three sectors: primary, secondary and tertiary. The primary sector relates to the transformation and the extraction of Raw Material (RM) from the land or sea (e.g. oil, iron ore, timber and fish). Some examples of industries within this sector are mining, quarrying, fishing, forestry, and farming. These materials are then used in industries from the secondary sector, which may also be called Manufacturing and Industry sector, or production sector, in which RM are transformed into finished goods on a large scale. This sector includes all branches of human activities that transform RM into products or goods, as secondary processing of RM, food Manufacturing, textile Manufacturing and Industry. Finally, the tertiary sector, also known as service sector, includes all branches of human activity whose essence is to provide services, thus contributing to physical/mechanical work, knowledge, financial resources, infrastructure, goods or a combination of those (Kenessey, 1987; Wolfe, 1955). This doctoral project takes places is a multinational company, founded in Portugal, and active in seven countries worldwide: Portugal, China, Ireland, Switzerland, USA, India and Japan. In this project the focus will be the factory in Portugal. The domain focus, where this company is positioned, is the secondary sector Industry or Manufacturing. With regard to the areas composing this specific context, this company’s industrial site is divided in three areas: Warehouse, Production and Laboratories. To have a better understanding about the current state of the organization and the relationship between those areas, Figure 2 presents the current, known as AS-IS, architecture and workflow of materials (RM, Samples) and Information between the different areas and the softwares that they use. With regard to the Warehouse area, this is the place where the products are received and shipped. The products that are provided by the suppliers are named RM. In the case of Intermediate (IN) Products and Final Product (FP), these arrive in the Warehouse from the Production area. The Production area is where the product manufacturing is performed. This area receives RM from the Warehouse and during the production process In Process Control (IPC) samples are created. The IPC sample is critical to the production process, and production may stop until the samples are approved by the Analytical Laboratories (AL). At the end of the process, IN and FP samples are created and the product packaging is sent to the Warehouse. With regard to AL, we can identify four main events, namely, planning the arrival of samples, the arrival of samples, weekly planning and scheduling, and testing. Both branches also require continuous support from Quality Control (QC) Laboratories. These Laboratories evaluate products throughout their whole life-cycle, assuring rule compliance to Good Manufacturing Practices (GMP), while complying with Good Laboratory Practices. This regulation is important to assure the pharmaceutical product’s safety and control, as well as its continued quality. With the recent company growth as a Contract Manufacturing 3
CHAPTER 1. INTRODUCTION Material Suppliers Raw Materials Analytical Material Warehouse RM, IN, IPC, FP Samples Intermediate & Final Products Production Notification of Acceptance Register and Stores Instruments Allocations Register and Stores Information about the Samples and analysis Material Requests Laboratories Raw and Analytical Materials Lead Time Notification of Production Start Production Orders Informations Materials needed in Production Material Requests and Arrivals ERP Database 2 Database 1 Physical Data Digital Data Figure 2: AS-IS Architecture of the Organization. Organization, the mix of products and laboratory tests has increased laboratory process complexity. Under this context, AL are fundamental to the company and to QC. In relation to the flow between the three areas (Warehouse, Production and Laboratories) and focusing more specifically on the Laboratories, it is important to point out that the Warehouse sends the physical samples of RM to the Production and here, an employee sends by email to the Laboratories, the names of the RM so that they can collect them. When the analyst is in charge of going to the Manufacturing building to raise the RM sample, it has to enter the aforementioned matter in the laboratory by registering it in “Database 2”. Then, since the RM sample is already in the laboratory, the analytical test can be carried out or, if it is not possible to start the analytical test at that moment, the RM is stored locally in the laboratory. Regarding the workflow between AL and the Production area, there are certain moments in production when it is necessary to take samples from the production line and send them to the AL to ensure that the production of the products is up to standards quality requirements for the products. The sample is taken by the Operator of the production machines in certain periods of time defined in the production sheet and then the analyst will collect the samples to analyze them in the AL. These samples are called IPC and have analysis priority over the remaining samples. The IPC Sample data is generated automatically in “Database 2” after the start of a Production Order. In the AL, the analysts usually do not record the sample arrival in the “Database 2” software, as they usually perform the analysis in the samples and when the analysis is done they write the sample 4
1.3. OBJECTIVES arrival and the test result at the same time in “Database 2”. This happens because in the analysts have to record the sample state and analysis procedure in physical documents named Logbooks. Regarding to the analytical tests that are carried out in the laboratory, if an analysis which is running does not get approval at the end due to deviations from the standard, the process is repeated to try to check where the error occurred. If the error occurs before the injection phase, there is no event registration. In case the analytical test goes wrong and this error occurs after the sample injection, it is communicated to QC, which will check if the problem comes from the components or materials that are being used and the deviation is reported in the Corrective Action and Preventive Action software. The AL have vast data stored in physical documents, which reduces the productivity during the analytical tests. This happens because the analyst must have to write all the products used and all the conditions of the analytical tests. However the AL use some Information Technology (IT) applications to help the management of the Laboratories. The application used in the Laboratories an Enterprise Resource Planning (ERP) for material requests and to have information about the Production Orders. Regarding the Samples management and tests, the Laboratories uses the “Database 2” to register the samples and the result of the tests performed to the samples and the “Database 1” to view the allocation of the HPLC and GC instruments. For privacy purposes, the names of the softwares and databases used in the AL are anonymized in this document. The implementation of an Intelligent Decision Support Systems (IDSS) can be potentially useful to create a ground truth of data by integrating the data from the different softwares used in the Laboratories. This would provide for the Analysts, new insights and improve various processes performed in the Laboratories. In this doctoral project, the goal is to improve the functioning of the AL by using Business Analytics techniques in the Industry 4.0 context, aiming to solve the analyzed company needs. 1.3 Objectives This PhD program, as stated before, was developed within an R&D project that aimed to implement three different WP in AL. In what concerns WP3, where this thesis is inserted, the objective was to know how Business Analytics techniques can improve processes in AL. Based on that, the research question to be answered in this PhD project is: How can an Intelligent Decision Support System (IDSS) be designed and implemented under the Industry 4.0 concept to create value in the Analytical Laboratories of a Chemical Industry? To answer the Research Question defined, we addressed the following intermediates objectives: • Conduct a Systematic Literature Review (SLR) on the use of Business Analytics techniques in Industry 4.0, to verify what type of techniques have been used in this context, which areas of industry are embracing the concept of Industry 4.0 and what are the open research gaps regarding the use of Business Analytics in Industry 4.0. 5
CHAPTER 1. INTRODUCTION • Develop predictive models that have the ability to predict the arrival of samples at the AL. These data-driven models will have to be able to predict the arrival of IPC samples at the Laboratories with high reliability in terms of arrival intervals, such that the analyst has time to prepare the materials in advance to analyze the samples in time, thus avoiding delays in the production process. Prediction models will have to be adaptable over time and be able to choose automatically the best algorithm for each training time interval. • Create ML models that are able to predict the consumption of materials in the AL based on historical consumption and the tests used. This algorithm will consume the forecasts of arrival of samples to the Laboratories (which contain the information about the analysis that will be performed) and will have to be able to timely forecast the requests of materials to the Warehouse, such that the laboratory does not have a shortage of materials for the analysis to be performed. If this happens, it may lead to delays in the analysis and, consequently, in the production process. • Develop models that are able to assign the best instrument for each analysis, taking into account the specific analysis, the maximum load of the instrument and the instrument’s capabilities, in order to make a more equitable distribution of the instruments to be assigned to the analyses. • Develop and evaluate an architecture for a IDSS that can be applied to all ALs in the Chemical Industry within an Industry 4.0 context. This architecture will encompass the aforementioned models, which must also comply with the current conditions in the Chemical Industry. The planned architecture contains a set of models that encompass all types of Gartner Analytics (Descriptive, Predictive, and Prescriptive). 1.4 Research Methodology In this PhD program, which is essentially a research project that involves the development of an artifact - an IDSS system - we use a Design Research, more specifically the Design Science Research Methodology for Information Systems (DSRM-IS), as our research methodology. This methodology consists of a set of techniques and methods with the objective of developing an IS artifact. Figure 3 presents the methodology and in it we can identify its five steps: Awareness of Problem, Suggestion, Development, Evaluation and Conclusion. In this section, each of these steps will be detailed when applied in the development of this PhD. Within the Chemical Industry, there are AL that are very important for the proper functioning of the industry because it is in these Laboratories that all products used and produced in the organization are analyzed to confirm that they are within the quality standards. However, currently much of the work that is done in the laboratory is manual and the communication that the laboratory makes with other entities is done manually, and delays in these processes can lead to pauses in production, which is not desirable. The first step of this project is the awareness of this issue, where a Systematic Literature Review was 6
1.4. RESEARCH METHODOLOGY Knowledge Contribution Knowledge Flows Process Steps Outputs Awareness of Problem Suggestion Development Evaluation Conclusion Design Science Knowlege Circumscription Proposal Tentative Design Artifact Performance Measures Results Figure 3: DSRM-IS Model, adapted from Vaishnavi and Kuechler (2004). performed and concluded that this situation in the AL can be improved with a set of novel techniques, along with new contributions that can be made in terms of business and body of knowledge. Therefore, the problem formulated arises that with the creation of an integrated system that uses different Business Analytics techniques (Descriptive, Predictive and Prescriptive), it will help the decision making process in Laboratories, as well as the streamlining of some of their processes. Furthermore, the exploration and implementation of these techniques can lead to scientific contributions within the subject of Business Analytics. The proposed solution was designed in the next step of the methodology. The next step of DSRM-IS aims at presenting the proposed solution to the problem formulated, where it will have to be implemented and evaluated in order to increase the body of existing knowledge and contribute to the resolution of the problem mentioned. As far as this PhD project is concerned, the proposed solution consists of an IDSS that uses Descriptive, Predictive and Prescriptive Business Analytics techniques to help decision making in AL. After the presentation of the proposed solution, the next step in this methodology is related to the development of the proposed solution. It is important to mention that the DSRM-IS methodology is cyclic, which means that the previous phases can be subject to change whenever such change is needed. For the development of this solution, it was divided into three phases, where the first phase was to create a hybrid architecture that uses Automated Machine Learning (AutoML) to predict the arrival of IPC samples to the ALs. The second phase of the development, focused on predicting the consumption of materials in the AL 7
CHAPTER 1. INTRODUCTION also using AutoML. Finally, the last phase, was the development of the method of allocating instruments to analyses to be performed and the creation of the IDSS dashboards. To develop this solution, R and Python programming languages were used. The output of this step is our IS artifact: an IDSS for AL in the Chemical Industry. Once an artifact is created, the DSRM-IS methodology assumes that it is evaluated by considering performance measures. For the predictive algorithms, to evaluate the performance, we used a wide range of performance metrics, namely Mean Absolute Error (MAE), Normalized Mean Absolute Error (NMAE), Mean Squared Error (MSE), Root Mean Square Error (RMSE), Regression Error Characteristic (REC), Area of Regression Error Characteristic (AREC) and R2. To evaluate the performance of our algorithms over time, we use a Rolling Window (RW) with a 20-week window to ensure the robustness of our algorithms. To evaluate our fully integrated IDSS, we uses questionnaires with 10 questions from Technology Acceptance Model (TAM) version 3 (Venkatesh & Bala, 2008), together with a matrix comparing the current functionalities (As-Is) with the functionalities provided by our IDSS. This evaluation is complemented with open interviews from the laboratory managers. When the results of the previous step are at a satisfactory level of acceptance, we move on to the last step of this methodology. In this doctoral project, this level of acceptance would mean that the system developed has a better performance, which is statistically proven when compared to existing methods and if the system reaches all the established objectives. Therefore, the conclusions will be withdrawn, whereby this usually means the publication of scientific publications. This project had as a result of its scientific process three conference papers and one journal article published. These articles are detailed in the next section. 1.5 Contributions This thesis includes a collection of research paper that were written during the execution of this PhD project with the goal of reaching the research objectives outlined in the previous section. Section 2.1 presents a SLR regarding the use of Business Analytics techniques in Industry 4.0 covering a selection of 169 papers obtained from six major scientific publication sources from 2010 to March 2020. The selected papers were first classified in three major types, namely, Practical Application, Reviews and Framework Proposal. Then, we analyzed with more detail the practical application studies, which were further divided into the three main categories of the Gartner analytical maturity model: Descriptive Analytics, Predictive Analytics and Prescriptive Analytics. In particular, we characterized the distinct analytics studies in terms of the Industry application and data context used, impact in terms of their Technology Readiness Level (TRL) and selected data modeling method. Our SLR analysis provides a mapping of how data-based Industry 4.0 expert systems are currently used, disclosing also research gaps and future research opportunities. This work resulted in a journal paper: • A. J. Silva, P. Cortez, C. Pereira, A. Pilastri, Business Analytics in Industry 4.0: A systematic 8
1.5. CONTRIBUTIONS review Expert Systems , e12741 (2021) DOI: https://doi.org/10.1111/exsy.12741 Section 3.2 addresses one of the major problems in the AL, the sample arrival, in this specific case, the IPC samples, as those have priority to be analyzed in order to avoid the stop of the production. The forecasting of sample arrival at the Laboratories is crucial for preparing the analytical materials on time, in order to analyze those samples at the Laboratories. To predict the sample arrival, different Cross-Industry Standard Process for Data Mining (CRISP-DM) iterations were performed, each focusing on a different regression approach. An AutoML was adopted during the modeling stage of CRISP-DM. Using recent realworld data from the Chemical organization, it was concluded that a proposed two-stage Machine Learning (ML) model was competitive and provided interesting predictions to support the laboratory management decisions (e.g.,preparation of testing instruments). This work was published in the following conference: • A. J. Silva, P. Cortez, A. Pilastri, Chemical Laboratories 4.0: A Two-stage Machine Learning System for Predicting the Arrival of Samples Artificial Intelligence Applications and Innovations , Springer International Publishing, 232-243 (2020) DOI: https://doi.org/10.1007/9783-030-49186-4_20 Section 3.3 focused also on the improvement of the AL management, more specifically the stock management, where the goal was to predict the material consumption in the AL based on the week plans of samples analysis using ML techniques. Several CRISP-DM iterations were performed and, to reduce the modeling effort, an AutoML was used to select the best ML model. Using real data from the Chemical company and a realistic rolling window evaluation, several ML train and test iterations were executed. The AutoML results were compared with two time series forecasting methods, the ARIMA methodology and a deep learning Long Short-Term Memory (LSTM) model. Overall, competitive results were achieved by the best AutoML models, particularly for the top 10 set of materials. This study resulted in the following conference paper: • A. J. Silva, P. Cortez An Automated Machine Learning Approach for Predicting Chemical Laboratory Material Consumption Artificial Intelligence Applications and Innovations , Springer International Publishing, 105-116 (2021) DOI: https://doi.org/10.1007/978-3-030-79150-6_9 Section 3.4 presents an IDSS to enhance the management of AL of a company operating in the Chemical Industry. This IDSS incorporates two predictive ML models, related with the prediction of the arrival of samples at the AL and the consumption of AL materials, which are then used to perform Prescriptive Analytics for AL instrument allocation tasks. The IDSS is also complemented with Descriptive Analytics of instrument similarities regarding the tests performed, for better supporting the AL manager decisions. The IDSS includes interactive dashboards and it was successfully validated by the AL managers using the TAM model 3 and open interviews, which resulted in a positive feedback. This work was accepted and therefore published in the following conference: 9
CHAPTER 1. INTRODUCTION • A. J. Silva, P. Cortez An Industry 4.0 Intelligent Decision Support System for Analytical Laboratories Artificial Intelligence Applications and Innovations , Springer International Publishing, 159-169 (2022) DOI: https://doi.org/10.1007/978-3-031-08337-2_14 1.6 Thesis Organization This thesis is structured as follows: • Chapter 1 describes the motivation, problem formulation, research objectives, contributions and PhD organization of this thesis. • Chapter 2 presents the main background associated with this PhD work. The first section details a SLR study that was performed on the topic of Business Analytics applied within the Industry 4.0 concept. The second section presents additional topics that were not discussed in the SLR but that are relevant for this PhD. Finally, the third section surveys studies that involve the usage of Business Analytics within the Chemical industry domain. • Chapter 3 presents the proposed methods and conducted experiments that led to the design of the proposed IDSS. In Section 3.2 we present the two-stage model to predict the arrival of samples at the AL. In Section 3.3 we present the model for predicting material consumption in the AL. And in Section 3.4 we detail the full designed IDSS, which includes a prescriptive model for instrument allocation as well as the development of the Dashboards used in the IDSS. • Finally, Chapter 4 summarizes the main conclusions of this PhD work, discussing some of its main impacts and limitations. Moreover, the last section presents future lines of research. 10
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Table 1: Literature surveys about the topic of Business Analytics in Industry 4.0. Reference Industry Sector Search Method a Descriptive Analytics Predictive Analytics Prescriptive Analytics O’Donovan et al. (2015) Manufacturing SLR X Chiang et al. (2017) Manufacturing Manual X Nikolic et al. (2017) Manufacturing AA X Uhlmann et al. (2017) Manufacturing Manual X X. Xu and Hua (2017) Manufacturing AA X X J. Yang et al. (2017) Manufacturing AA X Bordeleau et al. (2018) All Industry SLR X Sharp et al. (2018) Manufacturing AA X X Diez-Olivan et al. (2018) Manufacturing Manual X X X Qi and Tao (2018) Manufacturing Manual X Muhuri et al. (2019) All Industry AA X X Bakar et al. (2019) Manufacturing Manual X This Review All Industry SLR X X X a Automatic Analysis (AA), Systematic Literature Review (SLR) monitoring and Predictive Maintenance in the availability oriented business model. They also studied, based on practical examples, the organizational prerequisites for an implementation of these techniques in the industry. X. Xu and Hua (2017) summarized and analyzed the current research status for industrial Big Data Analysis in smart factories (both domestic and abroad). Also, they proposed research strategies for Industrial Big Data Analysis, including acquisition schemes, ontology modeling, predictive diagnostic methods based on Deep Neural Networks (DNN) and three-dimensional self-organized reconfiguration mechanism. In the area of Augmented Reality solutions, J. Yang et al. (2017) presented a comprehensive survey of AI in 3D painting to detect defective products in the Industry 4.0 context. The survey only analyzed Predictive Analytics techniques. Bordeleau et al. (2018) also performed a literature review of Business Intelligence in the context of Industry 4.0. The goal was to understand how Business Intelligence and data analysis generate value creation in manufacturing companies. This review only studied Descriptive Analytics. Sharp et al. (2018) presented another literature review about the use and development of Machine Learning in smart manufacturing. We note that this review studied practical cases that used Machine Learning in contexts different to the Industry 4.0 context. They reviewed the articles published between 2007 until 2017, while the Industry 4.0 concept was introduced in the 2010s. Moreover, the authors only analyzed the Diagnostic and Predictive Analytics. More recently, Diez-Olivan et al. (2018) presented a survey of the recent developments in data fusion and Machine Learning for industrial prognosis during the Industry 4.0 context. In the same year, Muhuri et al. (2019) performed a literature review about the growth of the Industry 4.0 in the last years. Bakar et al. (2019) presented a survey regarding the use of Metaheuristics techniques and Robotic Assembly Line Balancing in the Manufacturing industry. This SLR review is more focused on the whole Industry 4.0 concept, and thus it does not detail much the Business Analytics methods. A summary of the related work is presented in Table 1. None of the reviews analyzed addressed all main Gartner’s Analytical levels. In contrast, this SLR contains a stronger focus on the Descriptive, Predictive and Prescriptive analytics, when applied to the context of the Industry 4.0. Moreover, we 17
CHAPTER 2. BACKGROUND Table 2: Summary of the literature search protocol. Subject Business Analytics in Industry 4.0 Time period January 2010 to March 2020 Search Engines Scopus, ScienceDirect, SpringerLink, IEEE Xplore, Google Scholar, Google Books Search Criteria English; Title, abstract and keywords OR All (except full text) Search Query ”Industry 4.0 + Decision Support Systems”, ”Industry 4.0 + Business Analytics”, ”Industry 4.0 + Predictive Analytics”, ”Industry 4.0 + Machine Learning”, ”Industry 4.0 + Data Mining”, ”Industry 4.0 + Text Mining”, ”Industry 4.0 + Process Mining”, ”Industry 4.0 + Forecasting”, ”Industry 4.0 + Metaheuristic” particularly detail the practical applications, allowing to characterize the main business goals, data usage, modeling methods and obtained impacts. It should also be noted that most surveys consider only the Manufacturing sector, which is where the Industry 4.0 concept is producing a higher impact. Indeed, while this SLR considers all industry sectors, the selected practical research works in this SLR are highly related with the Manufacturing sector (as shown in Section 2.1.4.1). 2.1.3 Literature Review Method 2.1.3.1 Paper Collection We performed a manual SLR review, similar to what was proposed by Kitchenham et al. (2009). For this literature review, we used several scientific search engines, in order to search for the relevant documents: Google Scholar (https://scholar.google.com/), Google Books (https://books.google.com/), ScienceDirect (https://www.sciencedirect.com/), SpringerLink (https://link.springer.com/), Scopus (https: //www.scopus.com/home.uri) and IEEE Xplore (https://ieeexplore.ieee.org/Xplore/home.jsp). The term ”Industry 4.0”was coined in 2010. As shown in Figure 4, the Web interest in the term starts from 2010, although the popularity only increases substantially after 2014. Thus, we have retrieved articles that were published since 2010 until March 2020 (when this SLR was executed). Using the listed search engines, we performed several queries, using the combinations of the following keywords: ”Industry 4.0”, ”Decision Support Systems”, ”Business Analytics”, ”Predictive Analytics”, ”Machine Learning”, ”Data Mining”, ”Text Mining”, ”Process Mining”, ”Forecasting”, and ”Metaheuristic”. Table 2 presents the literature search protocol used during this SLR. Paper Selection Table 3 presents the distribution numbers of the collected scientific publications for the different search engines used. The paper search queries resulted in a total of 285 articles. All retrieved documents were manually inspected to check their relevance. First, the title and abstract was read. When the abstract was not conclusive, a more in-depth reading of the article was performed, in order to verify if the document fits the SLR goal. The manual inspection filtered 116 papers that were considered irrelevant 18
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Figure 4: Evolution of the interest in the term ”Industry 4.0”in Google Trends. Table 3: Distribution of papers obtained by each database. Database Quantity Scopus 202 ScienceDirect 75 SpringerLink 35 IEEE Xplore 60 Google Scholar 35 Google Books 5 Total with duplicates 390 Total without duplicates 285 19
CHAPTER 2. BACKGROUND for the survey, thus resulting in a total of 169 articles that were selected. 2.1.4 Literature Review Analysis 2.1.4.1 Quantitative Analysis As stated earlier, 169 papers were selected for this literature review. To make a general overview about the papers selected, a quantitative analysis was performed, in which the papers are characterized according to the year of publication and the paper type. Paper type The papers collected were manually inspected and divided into the three different categories proposed in Öchsner (2013): • Practical Application - These papers describe and discuss real implementation results of a framework, methodology, method or IT in one or more application domain areas; • Reviews - Articles of literature review (such as this SLR), with the main objective of performing a survey of the state-of-the-art on a certain scientific research topic area, possibly identifying research gaps; and • Framework Proposal - The aim is to document the proposal of a new framework developed by the authors. However, these articles do not have a specific application target, thus the authors do not validate the framework in a real-world environment. Table 4 shows the respective distribution of the selected 169 papers in terms of the three main paper categories. The majority of the selected papers are Practical Application ones (139 papers). There are 11 Table 4: Distribution of the three main paper types. Paper Type Quantity Practical Application 139 Reviews 12 Framework Proposal 18 Total 169 papers that were categorized as Reviews and 18 publications categorized as Framework Proposal. Given that this survey is more focused on practical usage of Business Analytics, we will only further detail and analyze the 139 Practical Application studies. The quantitative analysis includes the industry sector, the Gartner Analytic type and year, and finally the paper keyword frequencies. 20
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Table 5: Distribution of the Practical Applications per industry sector. Industry Sector Quantity Manufacturing 130 Transportation and Warehousing and Utilities 3 Construction 2 Educational Services, and Health Care and Social Assistance 1 Agriculture, Forestry, Fishing, and Hunting, and Mining 2 Finance and Insurance, and Real Estate, and Rental and Leasing 1 Total 139 Industry sectors of the Practical Applications To describe the Industry sections we adopted the Standard Industrial Classification Bureau, 2017, which includes five main categories listed in Table 5. The Manufacturing sector is by a large margin the sector with most Industry 4.0 practical applications of Business Analytics, with 130 papers. This happens because the manufacturing sector is a vast sector that includes a relevant number of production processes, widely used by several industries. The manufacturing sector has also high Business Analytics needs. For instance, the shop floor usually has different kinds of machines, which should work efficiently and produce quality products. Thus, Predictive Maintenance and automatic quality inspection/prediction methods, based on data-driven models, can be used to enhance the manufacturing process. The other industry sectors have much less practical application works. Within the Transportation and Warehousing and Utilities sector, the surveyed papers relate with three practical applications. In the Eolic Energy area, Canizo et al. (2017) presented a data-driven solution deployed in a cloud that used Random Forest (RF) for predicting failures on wind turbines. In the transformation energy field, Bagheri et al. (2018) analyzed the analytical approach to the transformer vibration modeling, using Machine Learning techniques such as Linear Regression (LinR), Model Trees, Support Vector Regression with Gaussian Kernel and Multilayer Perceptron, and also signal techniques to develop prognosis models of transformer operating condition based on vibration signals. Masoudinejad et al. (2018) proposed a set of Support Vector Machine (SVM) algorithms, addressing indoor localization within a warehouse. The Construction sector has two practical applications. J. Lee et al. (2014) made a review about the trend of the manufacturing service transformation in Big Data and proposed a framework for sustainable innovative service. The data used to make the case study came from sensors installed in a bulldozer. They used a Bayesian Belief Network to classify if the engine had some problem or malfunction and used a Fuzzy-Logic based algorithm to predict the remaining useful life of the engine. R. Costa et al. (2017) proposed a system with the aim to create knowledge representations from unstructured data sources used in a construction environment, based on enriched semantic vectors. Regarding the Educational Services, and Health Care and Social Assistance sector, Bordel and Alcarria (2017) presented a solution to automatically assess the human motivation in Industry 4.0 scenarios with the use of an ambient intelligence infrastructure. Turning to the Agriculture, Forestry, Fishing, and Hunting, 21
CHAPTER 2. BACKGROUND Table 6: Distribution of the Practical Applications for the three Analytics types and year of publication. Year Descriptive Predictive Prescriptive Total Analytics Analytics Analytics 2015 2 0 0 2 2016 2 8 2 12 2017 8 17 2 27 2018 9 21 5 35 2019 2 25 10 37 2020 0 9 8 17 Total 23 80 27 130 and Mining sector, Teschemacher and Reinhart (2017) used Ant-Colony Optimization algorithms to enable dynamic milk-run logistics. Also, Dutta et al. (2018) implemented a Machine Learning based interactive architecture for industrial scale prediction for dynamic distribution of water resources across the continent and, at the same time, keeping four corners of Industry 4.0 in place. The algorithms tested were LinR, Bayesian Ridge Regression, Logistic Regression (LogR), Linear Discriminant Analysis, Adaptive NeuroFuzzy Inference System, Multi-Layer Perceptron, and Radial Basis Function Network. Finally, within the Finance and Insurance, and Real Estate, and Rental and Leasing sector, Ma and Li (2018) used a Grey Model to predict eight indexes of the tertiary industry. 2.1.4.2 Analytics Type Table 6 shows the distribution of the selected Practical Application papers in terms of publication year and analytics type. The most common type is the Predictive Analytics level, with 80 applications, followed by the Prescriptive Analytics, with 27 applications, while the Descriptive Analytics were only addressed in 23 applications. The smaller number associated with the Prescriptive and Descriptive Analytics denote an important research gap. The lack of further Prescriptive studies is probably due to two main reasons. Firstly, the Industry 4.0 concept implementation is very recent (just a few years). Most of its initial implementation effort is devoted to setting the right infrastructure to generate and collect data, and Business Analytics can only be applied after collecting enough historical data. Secondly, Prescriptive Analytics are more complex than other types of data analyses (Koch, 2015). As more mature Industry 4.0 applications are implemented, we expect this gap to be reduced. It is also interesting to note that there are more Predictive Application studies than Descriptive ones. This behavior might be explained by the current Machine Learning hype. Also, building a stable and valuable Data Warehousing system, which results in better Descriptive analysis, requires several Extract, Transform, Load (ETL) processes that are often costly, requiring manual effort and time, but that do not tend to translate into novel methodologies or interesting application usages that justify a research publication. Overall, the yearly numbers from Table 6 show a substantial growth in the number of publications starting from 2017: 27 papers in 2017; 35 works in 2018; and 37 research publications in 2019 (the 8 papers from the year of 2020 report only until the 22
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Figure 5: Word cloud of the keywords (left) and top 10 term frequency values (right). month of March). Keywords frequencies The last quantitative analysis is obtained by applying a word cloud technique to the 112 application paper keywords. We have selected keywords because these help to index and classify papers, facilitating research queries. The word cloud analysis was performed using R tool with the package wordcloud. The word cloud is presented in Figure 5, which also details the top term frequency numeric values. The most frequent term is ”Industry”, followed by ”data”, ”learning”and ”manufacturing”. Other terms such as “maintenance”, ”machine”and ”predictive” are also popular, which aligns with Table 6, since most practical applications use Predictive Analytics. 2.1.4.3 Qualitative Analysis The qualitative analysis was executed by a manual inspection of the selected practical papers. The description of practical cases are divided by the analytics type (Descriptive, Predictive and Prescriptive), using a chronological order. Each practical application is briefly described, including the: •Function – Industry 4.0 function area, which is categorized by the four main functions of the Industry 4.0 architecture presented by Qin et al. (2016): Hardware Connection (HC), focuses on hardware development (e.g., sensor network); Information Discovery (ID), where the raw data is transformed into useful knowledge; Predictive Maintenance (PdM), aiming to anticipate maintenance issues; and Intelligent Production (IP), automating or adapting the production process. •Data – type of industry data used (e.g., generated by a production machine, captured image). 23
CHAPTER 2. BACKGROUND Table 7: Description of the Technology Readiness Levels (TRL). Phase Level Definition Research TRL 1 Basic research TRL 2 Technology formulation TRL 3 Concept validation Development TRL 4 Prototype in laboratory environment TRL 5 Prototype in relevant environment TRL 6 Prototype system tested in relevant environment Deployment TRL 7 Demonstration system in operational pre-commercial environment TRL 8 First commercial system, ready for operational environment TRL 9 Full commercial system with general availability •Sector – addressed industry sector (e.g., aerospacial, automotive). •Goal – brief description of the application goal. •Impact – measured using the TRL scale, from 1 to 9 (Table 7) (ESRTC, 2009). •Modeling – Business Analytics method used to analyse the data. Table 8: Overview of the Practical Articles that used Descriptive Analytics Techniques Reference Func.1Data2Sector3Goal Impact Modeling4 Neuböck and Schrefl (2015) ID Pr ND New analysis graphs are proposed for building production insights (e.g., show urgent missing materials). 7 DW, AG Niño et al. (2015) IF MF CE Big Data Analytics for pursuing a servitization strategy. 2 DA 1Hardware Connection (HC), Information Discovery (ID), Intelligent Production (IP), Predictive Maintenance (PdM) 2Car Specification (CS), Grippers (G), Historical (H), Machine (MC), Manufacturing (MF), Production (Pr), Sensor (S), Sparse Data (SD), Temporal Logs (TL) 3Additive Manufacturing (AM), Aerospace (As), Automotive (A), Capital Equipment (CE), Chemical Industry (CI), Glass Industry (GI), Not Disclosed (ND), Semiconductor (SC), Spring Manufacturing (SM) 4Analysis Graph (AG), Artificial Neural Networks (ANN), Augmented Reality (AR), Back-Propagation Artificial Neural Networks (BPANN), Back-Propagation Neural Networks (BPNN), Browns Double Exponential Smoothing (BDES), Classification Trees (CT), Clustering (Cl), Cross-Departmental Data Analytics (CDDA), Data Analysis (DA), Data Warehouse (DW), Decision Trees (DT), Deep Learning (DL), Descriptor Silhouette (DS), Digital Twin (DigT), Failure Mode Metrics (FMM), Fuzzy Logic (FL), Genetic Algorithm (GA), Interpolation Fitting (IF), K-Mean Clustering (KMC), Linear Regression (LinR), Monkey Algorithm (MA), Neural Networks (NN), NeuroEndocrine-Inspired Manufacturing System (NEIMS), Partial Least Square (PLS), Residual Prediction Calculator (RPC), Self Organizing Map (SOM), Simulation (Sim), Standard Silhouette (SS), Two-Stage Clustering (TSC) 24
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Y.-M. Lee et al. (2016) ID S A Real-time analysis to explore the reasons for abnormality of load rate data of main shaft machine. 5 BPANN, TSC Tang et al. (2016) HC MF ND Intelligent architecture for the smart shop floor. 5 NEIMS Durakbasa et al. (2017) ID S ND Improve the quality of the manufacturing process. 2 FL Kirchen et al. (2017) ID S CI Explore signal data quality. 4 DA C.-J. Kuo et al. (2017) IP S SM Explore inexpensive add-on triaxial sensors for the monitoring of machinery. 5 NN Qin et al. (2017) ID MC AM Facilitate a better understanding of the energy consumption of digital production processes. 5 LinR, DT, BPNN Sanz et al. (2017) ID S A Advanced monitoring of an industrial process that integrates several data sources. 3 BDES Trunzer et al. (2017) ID S ND Classify failures in control valves. 4 FMM, GA Y. Wang et al. (2017) ID SD ND Methodology to enrich sparse data by fast and frugal reduced models. 3 Cl, CT Zheng and Wu (2017) ID Pr SC Smart spare parts inventory management system for semiconductors. 5 DA, Sim Birglen and Schlicht (2018) ID G A Review the characteristics of pneumatic, parallel, two-finger and industrial grippers. 3 DA Lenz et al. (2018) ID S ND Holistic approach for machine data analytics. 2 CDDA C. Lin and Yang (2018) HC S ND Intelligent Computing System to connect the different facilities in a logistic center. 6 MA, GA 25
CHAPTER 2. BACKGROUND Mozgova et al. (2018) ID S A Monitor actual stress state of a structural component and estimate its residual fatigue life. 6 RPC Ploennigs et al. (2018) ID, HC S ND Cognitive IoT architecture with scalability and self-learning capabilities. 5 AR Stürmlinger et al. (2018) HC S ND Development of a new generation of a manufacturing system. 5 DA Subakti and Jiang (2018) ID MC ND Augmented reality system to visualize and interact with machines in smart factories. 7 DL Tieng et al. (2018) ID S As Virtual metrology system for sampling. 4 BPNN, PLS, GA, IF Vathoopan et al. (2018) HC H ND Corrective maintenance using the digital twin of an automation model. 3 DigT Kaupp et al. (2019) IP TL GI Outlier identification to measure the glass quality. 5 NN Ventura et al. (2019) ID S, P ND Automatic industrial equipment maintenance system. 6 DS, SS, KMC Descriptive Analytics Table 8 presents an overview of the practical applications that used Descriptive Analytics techniques. As shown in the table, there is a diversity of Descriptive applications and adopted types of historical analyses. For instance, some studies perform a simple statistical analysis (Birglen & Schlicht, 2018; Lenz et al., 2018; Mozgova et al., 2018; Niño et al., 2015; Sanz et al., 2017; Stürmlinger et al., 2018; Tang et al., 2016; Ventura et al., 2019), while others use more sophisticated outlier detection (Y.-M. Lee et al., 2016; Trunzer et al., 2017) and clustering methods (Y. Wang et al., 2017). Some studies use data warehousing databases and dashboards (Kirchen et al., 2017; Neuböck & Schrefl, 2015; Vathoopan et al., 2018; Zheng & Wu, 2017), and other studies used Neural Networks (Kaupp et al., 2019; C.-J. Kuo et al., 2017; Qin et al., 2017; Subakti & Jiang, 2018; Tieng et al., 2018). Predictive Analytics The practical applications that used Predictive Analytics techniques are shown in Table 9. Predictive Analytics involve a set of data-driven models that are typically obtained by applying supervised Machine Learning algorithms. 26
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW W. J. Lee et al. (2019) PdM Pr ND Predictive Maintenance to monitor two machine tool system elements, the cutting tool, and the spindle motor. 5 SVM, DL Liulys (2019) PdM S El Open-source software to develop predictive maintenance applications with basic programming knowledge. 3 GBM, NN Massaro, Manfredonia, Galiano, and Xhahysa (2019) IP Pr Fu ANN to predict the product defects in a kitchen manufacturing Industry. 5 ANN Massaro, Manfredonia, Galiano, Pellicani, et al. (2019) IP S Fo Predict the humidity during the pasta production. 5 ANN Martinek and Krammer (2019) IP I El Machine Learning based prediction methods to optimize the process parameters of pin-in-paste. 5 ANN, ANFIS, GBDT Packianather et al. (2019) PdM ChL Hc Three phase methodology to automate quality control in healthcare clinical laboratory. 5 KNN Pinto and Cerquitelli (2019) PdM S Rb Predict the fault detection and remaining life estimation of robots. 5 SA, ERT, KNN, CNN Plehiers et al. (2019) IP S Ch Framework for chemical production in process-steam cracking to optimize the process control. 4 ANN 33
CHAPTER 2. BACKGROUND Proto et al. (2019) IP S Ch PREdictive Maintenance service for Industrial procesSES (PREMISES) to predict alarms in slowly-degrading multicycle industrial process. 6 GBTC, RF Rogier and Mohamudally (2019) IP SolP En NN to predict the conversion of solar energy by a photovoltaic unit. 5 NN Rosli et al. (2019) PdM S SC Preventive maintenance for air booster compressor motor failure. 4 ANN, PSO Rousopoulou et al. (2019) PdM Pr Hc Predictive analytics for industrial ovens in the healthcare industry. 5 SVM Sellami et al. (2019) PdM Mc SC Predict machine failures and presented an algorithm for frequent chronicles extraction. 4 ClaspCPM Soto et al. (2019) PdM S ND IoT Machine Learning and orchestration to failure detection of surface mount devices during production. 4 NN, RF, GB Naskos et al. (2019) PdM Mc O Predictive Maintenance with applied unsupervised Machine Learning techniques to detect early oil leaks. 5 MCCOD Zenisek et al. (2019) PdM S ND Machine Learning algorithms to detect changing behavior to enhance the maintenance on a microscopic level. 3 RF, SVM, GPBSR T. Zhang et al. (2019) IP Pr El Random-SVM (R-SVM) to predict the quality of the TFT-LCD liquid. 4 RSVM Alasali et al. (2020) IP Mc Tr Predict the stochastic loads to improve the performance of a low voltage network. 6 MPC, SMPC Calabrese et al. (2020) PdM Mc Fu Machine Learning to predict the health status of a woodworking industrial machine. 6 GB, RF, EGB Q. Cao et al. (2020) PdM Pr ND Rule-based refinement approach for detect and predict anomalies. 4 RB 34
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Essien and Giannetti (2020) IP Mc ND Deep Learning model for univariate, multi-step machine speed forecasting in a manufacturing process. 4 DL Kabugo et al. (2020) IP Pr En Predict syngast heating value and hot flue gas temperature from data obtained from soft sensors. 5 NN Karakose and Yaman (2020) PdM S Tr Fuzzy system-based approach for Predictive Maintenance on electric railways. 4 CF Kim et al. (2020) IP Pr ND Predict the state of an unseen camera lens module using semi-supervised regression. 5 DL RuizSarmiento et al. (2020) PdM Pr SP Estimate and predict the gradual degradation of production machines. 5 BF de Sá et al. (2020) HC Pr ND Metaheuristics to identify data injection attacks by man-in-the-middle. 4 BSOA, GN, NII Predictive Analytics are the most used techniques in the practical applications obtained for this SLR. For instance, some studies perform classification techniques (Q. Cao et al., 2020; Kiangala & Wang, 2018; S. C. Li et al., 2017; Miškuf & Zolotová, 2016; Sellami et al., 2019), while other used regression techniques (Calabrese et al., 2020; Charest et al., 2018; Peralta et al., 2017; Rousopoulou et al., 2019). Simple Neural Networks (NN) are used in several research works such as (Cisotto & Herzallah, 2018; Kabugo et al., 2020; Miškuf & Zolotová, 2016; Soto et al., 2019; Spendla et al., 2017). Other studies used more advanced Deep Learning (DL) NN (Choi et al., 2017; Essien & Giannetti, 2020; H. Kuo & Faricha, 2016; W. J. Lee et al., 2019; Maggipinto et al., 2018). Furthermore, some of the surveyed Preditive Analytics used optimization techniques (e.g., Genetic Algorithm, Paticle Swarm Optimization) (Rosli et al., 2019; Saldivar, Goh, Li, Chen, et al., 2016; Saldivar, Goh, Li, Yu, et al., 2016), while other works focused on outliers detection and statistical analysis (Albers et al., 2017; Stein et al., 2016). Prescriptive Analytics The last table of this SLR (Table 10) presents the practical cases that used Prescriptive Analytics. These types of analytics aims to describe what courses of action may be taken in the future to optimize business processes in order to achieve business objectives. Typically, this is 35
CHAPTER 2. BACKGROUND achieved by associating decision alternatives (or choices) with estimated business outcomes. A diverse set of modeling tools can be used to obtain such analytics, namely optimization and simulation, design experimentation and scenario scheduling (Banerjee et al., 2013; Jugulum, 2016). The majority of the surveyed studies used optimization techniques. In particular, the most explored method was the Genetic Algorithm (Khayyam et al., 2019; D. Silva et al., 2020). Other authors (Ansari et al., 2019; Brik et al., 2019; Fu et al., 2018; H. Li, 2016; Qu et al., 2016; Tsourma et al., 2018; Uriarte et al., 2018), employed other optimization techniques, such as (Tsourma et al., 2018) that proposed a Task Distribution Engine to automate and optimize the task scheduling and resources assignment procedure in industrial environments. We also found studies that performed Prescriptive Analytics by using predictive models to directly perform actions: DL (Richter et al., 2017); Regression Trees and Nearest Neighbors (Romeo et al., 2018); and SVM combined with Q-Learning (Qu et al., 2016). Table 10: Overview of the Practical Articles that used Prescriptive Analytics Techniques Reference Func.9Data10 Sector11 Goal Impact Modeling12 H. Li (2016) IP Co ND Classification algorithm and Q-learning algorithm to reduce the electricity consumption in an automation system. 4 SVM, QL Qu et al. (2016) IP Pr ND Synchronized, station-based flow shop with multi-skill workforce and multiple types of machines. 3 RL, MARL, Op Klement and Silva (2017) IP Pr Pl Hybrid approach with List Algorithm and Metaheuristic to optimize planning, assignment, scheduling and lot sizing. 3 LA, SA Richter et al. (2017) IP Mc El Optimization techniques for the manufacturers and users of AOI machines. 2 DL Bányai et al. (2018) HC Ge Tr Black Hole Optimization for first mile and last mile supply. 6 BHO 9Hardware Connection (HC), Information Discovery (ID), Intelligent Production (IP), Predictive Maintenance (PdM) 10Conveyor (Co), Geospatial (Ge), Industrial (In), Machine (Mc), Network (N), Production (Pr), Sensor (S) 11Automotive (A), Chemical (Ch), Electronic (El), Lean (Le), Mechanical (MC), Not Disclosed (ND), Polymer (Pl), Transportation (Tr) 12Artificial Neural Networks (ANN), Black Hole Optimization (BHO), Constrained Optimization (CO), Coyote Optimization Algorithm (COA), Crow Search Algorithm (CSA), Decision Trees (DT), Deep Learning (DL), Fireworks Algorithm (FA), Fog Computing (FC), Genetic Algorithm (GA), Global Cheapest Arc (GCA), Grey Wolf Optimizer (GWO), Guided Local Search (GLS), Iterative Local Search (ILS), K-Nearest Neighbor (KNN), List Algorithm (LA), Memetic Algorithm (MmA), Mixed Integer Linear Programming Model (MILPM), Multi-Agent Reinforcement Learning (MARL), Multiple-layer perceptron neural network (MLPNN), Multi-Objective Optimization (MOO), Neighborhood Component Feature Selection (NCFS), Optimization (Op), Particle Swarm Optimization (PSO), Path Cheapest Arc Savings (PCAS), Prescriptive Maintenance Model (PriMa), Q-Learning (QL), Random Forest (RF), Regression Trees (RT), Reinforcement Learning (RL), Self Organizing Migrating Algorithm (SOMA), Simplified Swarn Optimization (SSO), Simulated Annealing (SA), Simulated Annealing Tabu Search (SATS), Simulation-based Multi-Objective Optimization (SBMOO), Support Vector Machines (SVM), Tabu Search (TbS), Variable Neighborhood Descent Based (VNDB), Variable Neighborhood Search (VNS), Whale Optimization Algorithm (WOA) 36
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Fu et al. (2018) IP In ND Two-objective stochastic flow-shop deteriorating and learning scheduling problem for advanced intelligent machines. 4 MOO, FA Romeo et al. (2018) IP Mc El Design Support System (DesSS) for the prediction and estimation of machine specification data. 4 DT, RT, KNN, NCFS Tsourma et al. (2018) IP In ND Task Distribution Engine to automate and optimize the task scheduling and resources assignment procedure in industrial environments. 5 CO Uriarte et al. (2018) IP Le ND Simulation and optimization to improve the lean efficiency, speeding up system improvements and reconfiguration. 2 SBMOO Ansari et al. (2019) IP Mc MD Prescriptive Maintenance model for production CPS. 6 PriMa Brik et al. (2019) IP In ND Fog computing architecture to deal with system disruption monitoring. 4 FC Khayyam et al. (2019) IP Pr Pl Genetic Algorithm to predict the stabilization process of a Plyacrylonitrile fiber structure. 5 GA Leite et al. (2019) IP Pr Pl Optimize the integrated planning and scheduling using Metaheuristic approach. 4 VNDB Liang et al. (2019) PdM S ND Memetic Algorithm and Variable Neighborhood Search to improve Predictive Maintenance. 4 MmA, VNS Negri et al. (2019) IP S ND Metaheuristics with Digital Twin for scheduling optimizations based on the equipment health predictions. 6 GA Pane et al. (2019) IP Mc MC Two reinforcement learning based compensation methods for robot manipulators. 5 RL Pierezan et al. (2019) IP S En Coyote Optimization Algorithm to optimize a heavy duty gas turbine used in power generation. 6 COA 37
CHAPTER 2. BACKGROUND Senkerik et al. (2019) IP S Ch Ensemble of strategies and Metaheuristic for optimization of waste processing batch reactor geometry and control 4 SOMA Yeh et al. (2019) IP N ND Optimization techniques to find the cost minimization deployment of a smart factory. 4 SSO Abdelmaguid (2020) IP Pr ND Algorithm to obtain optimal solutions for Dynamic Open Shop Scheduling Problem. 4 MILPM Abdirad et al. (2020) HC Ge A Two-stage metaheuristic to solve dynamic vehicle routing problem. 4 PCAS, GCA, GLS, SATS Abdous et al. (2020) IP S A Design semi-automated assembly lines using Machine Learning and Optimization techniques. 4 ILS Kharwar et al. (2020) IP S Pl Particle Swarn Optimization to optimize milling parameters (weight, spindle speed, feed rate and depth of cut). 4 PSO Y. Li et al. (2020) IP Pr Pl Hybrid model using Optimization and Machine Learning for production rescheduling. 4 GA, TbS, RF, SVM, MLPNN Milošević et al., 2020 IP Pr ND Compared three optimization algorithms for intelligent process planning optimization. 4 GWO, WOA, CSA Rahman et al. (2020) IP Pr ND PSO for line balancing and automated guided vehicles scheduling for smart assembly systems. 4 PSO D. Silva et al. (2020) IP Pr ND Hybrid ANN model and use GA for the multi-objective strength optimization of concrete with fiber. 5 ANN, GA 2.1.5 Discussion Figure 6 presents the Literature Map resulted from this SLR. This Literature Map contains three different levels of interactions, where the first level is the Analytics Level and the second level contains the components of the different Analytics application levels (Data Visualization, Detect Production Anomalies, 38
2.1. BUSINESS ANALYTICS IN INDUSTRY 4.0: A SYSTEMATIC LITERATURE REVIEW Improve Product Quality, Detect Costumers Needs, Predictive Maintenance and Resources Optimization). The last level presents the different techniques used for each component, as well as some studies that use these techniques. To simplify the visualization, the map only details business analytics techniques that were used in two or more practical cases. It is clear in Figure 6 that Supervised Learning techniques (Classification and Regression algorithms) are a popular approach of Business Analytics in Industry 4.0, being adopted in all the application types identified in this SLR. Statistical Data Analysis is a technique used mainly for Data Visualization, but it was also used for Predictive Maintenance (Mozgova et al., 2018), to Detect Anomalies in Production (Zheng & Wu, 2017) and to Improve Product Quality (Kirchen et al., 2017). Clustering is a more advanced technique compared to Statistical Data analysis, and is used to find Production Anomalities (Y. Wang et al., 2017), to improve the products quality (T. Lin et al., 2016), to detect costumers needs (Saldivar, Goh, Li, Yu, et al., 2016), and for predictive maintenance (Candanedo et al., 2019). Reinforcement Learning was used mostly for Resources Optimization (Pane et al., 2019; Qu et al., 2016), while Optimization techniques were used for Resources Optimization (Uriarte et al., 2018), to Detect Production Anomalies (Trunzer et al., 2017) and to Improve Product Quality (Khayyam et al., 2019). Regarding the Supervised Learning techniques, based on Classification and Regression algorithms, it is important to mention the popularity of NN (in their Artificial Neural Networks (ANN), DL, or Convolutional Neural Networks (CNN) forms), in the different Industry 4.0 areas. In effect, the use of NN reaches every area of application studied in this SLR with a total of 38 practical applications retrieved in this study. Moreover, the use of NN is growing over the time, with 4 applications in 2016, 10 in 2017, 7 in 2018, 14 in 2019, and 3 applications in the first months of 2020. The Literature Map from Figure 6 provides a general overview of the different application areas of Business Analytics in Industry 4.0, where it is clear that the areas of Improve Product Quality, Anomalies Detection and Predictive Maintenance are the most popular. While Business Analytics techniques can also be employed to optimize resources in the Industry or to Detect Costumers Needs, a small number of research application studies have addressed these topics, with 9 applications focused on Resources Optimization and 4 applications in Detect Costumer Needs. 2.1.6 Conclusions and research implications This section presents the results of this SLR to analyze the evolution and the application of Business Analytics techniques in the Industry 4.0 context. As stated in Section 2.1.1, the Research Question targeted by this SLR research is: How and in what areas of the industry are Business Analytics techniques being used in an Industry 4.0 context? The papers were surveyed by performing an initial keywords query on scientific search engines. Then, the retrieved papers were manually inspected by performing a careful analysis, to assure that the most relevant studies for this SLR were selected. Next, we have analyzed the selected papers in terms of both quantitative and qualitative elements. The quantitative analysis 39
CHAPTER 2. BACKGROUND Figure 6: Literature Map. showed that the most published type of paper is the Practical Application. As for the quantitative analysis, it consisted in a characterization of the Descriptive, Predictive and Prescriptive analytics in terms of what types of applications are implemented in the Industry, what are the techniques used in the practical applications and the impact of the results achieved. Considering the presented SLR we highlight that: • The application of Business Analytics techniques within the Industry 4.0 concept has grown in recent years and its popularity is still rising (as shown in Figure 4 and Table 6). Thus, there is a research opportunity for publishing more papers regarding Business Analytics applied to the Industry 4.0. • Manufacturing is the industry sector with the most practical applications (Table 5). One contributing factor for this phenomenon is that there has been a financial support for the adoption of innovative manufacturing techniques (European Commission, 2013). Nevertheless, there is a research opportunity set in terms of addressing other industry sectors, such as Transportation or Construction. • Within the manufacturing sector, most of the practical applications focused on problems existing in production lines, with different goals, such as detecting faults in production components, defective products, until monitoring the production process and optimization of productive components such as energy consumption and resources allocation. • Regarding the type of analytics, Descriptive Analytics involved a total of 23 practical applications, Predictive Analytics with 80 application studies and Prescriptive Analytics included 24 research 40
2.2. OTHER RELEVANT CONCEPTS works. The popularity of Predictive Analytics is being linked with the growing interest in the fields of Machine Learning and Data Science in the decade of 2010 (C. Costa & Santos, 2017). • Regarding the modeling techniques used, Supervised Learning was the most used approach, with NN being used in 39 applications, RF in 10 applications, SVM in 6 applications, Decision Tree in 5 applications and Rule-Based in 2 applications. Classical Statistical Data Analysis was used in 11 applications, Clustering was addressed in 6 applications, the same number as Optimization techniques, and Reinforcement Learning was employed in 2 applications. Given the current success of the DL field (Goodfellow et al., 2016), it is expected that the number of Industry 4.0 research works that use NN will further increase in the future. • Practical applications that use Descriptive Analytics are focused on analyzing the data obtained in order to find answers for diverse problems, such as verifying the tool wear through the time or what is the most common cause that leads to the equipment failure. • The practical applications that used Predictive Analytics were more focused in Predictive Maintenance, such as predict when the equipment will fail, or verify if the equipment is not corresponding in terms of its typical performance. • Practical applications that use Prescriptive Analytics target more on resources optimization, such as optimize the energy consumption or optimize the resources scheduling. However, the SLR results reveal that there is still a scarce number of research studies that use Prescriptive Analytics techniques within the Industry 4.0. Therefore, there is a huge potential for future research on more Prescriptive Analytics studies since there is a large number of industrial needs that are related with resource optimization and schedulingMoreover, as pointed out by Davenport (2013), these are the analytics “that tell you what to do” and thus hold a higher business value by providing an actionable knowledge for the industry. Thus, in future works, we believe there will be an increase of Prescriptive Analytics applications for the Industry 4.0. This SLR reviewed research papers published in the last decade (from 2010 to 2020). In the next decade, it is expected that Business Analytics will be more prevalent in the Industry, due to further advances in AI and ML. In particular, as the European Commission plans a future investment of 7.5 Billion EUR in the areas of Advanced Computing and AI (Commission, 2020), several of these funds will be devoted to Industry applications, which surely will be reflected in an increased number of research papers. 2.2 Other Relevant Concepts In this section we present theoretical concepts that were not addressed in the SLR, but are relevant for this PhD work, as well as a survey of the state-of-the-art related to the aplication of Business Analytics in the Chemical domain and its AL. 41
CHAPTER 2. BACKGROUND 2.2.1 Machine Learning (ML) ML is a branch of AI that enables computer algorithms to learn from experience without explicitly being programmed (Breiman, 2001; Obermeyer & Emanuel, 2016). ML uses computers with the objective of simulating human learning and allows the machines to identify and acquire knowledge from the real world, and improve the performance of tasks with the knowledge obtained (Portugal et al., 2018). Mitchell (1997) defined ML as ”a computer program is said to learn from experience E with respect to some class of tasks T and performance measure P, if its performance at tasks in T, as measured by P, improves with experience E”. There are four main types of methods in ML: supervised learning, unsupervised learning, semi-supervised learning and reinforcement learning (Portugal et al., 2018). This PhD thesis is mostly focused on Supervised Learning, where the learning algorithms have labeled training data, meaning that each input example contains a target output. The task is then to learn an implicit mapping function that exists in the training data. Such function should be able to generalize to new situations (Portugal et al., 2018). Therefore, supervised learning plays a key role in predictive analytics. The two main supervised learning types are Classification and Regression (Zhu et al., 2003). Common Classification methods include Decision Trees (DT), K-Nearest Neighbor (KNN), SVM, NN and RF (Larose, 2004). A few examples of pure Regression methods are LinR, Lasso and Elastic nets (Ogutu et al., 2012). SVM, RF and NN can also be applied to regression (Rodriguez-Galiano et al., 2015; Steyerberg et al., 2014). Recently, there has been an increasing interest in the adoption of DL NN, which can be applied to both classification and regression, and that achieved competitive results in several ML challenges (e.g., object recognition from images) (Goodfellow et al., 2016). Unsupervised learning does not have a target variable defined, it is the algorithm that searches for patterns and structures the information among data variables. One of the most used unsupervised techniques is clustering. A clustering algorithm automatically groups data variables or examples according to a distance function and clustering metric. Another popular unsupervised learning approach is rule association mining (Larose, 2004). These algorithms can be used to obtain a descriptive knowledge. The semi-supervised approach falls between the supervised and unsupervised approaches. It often assumes that there some labeled data and also unlabeled data. The semi-supervised approach is useful when labeling data is costly to obtain, such as requiring a manual effort (Zhu et al., 2003). Active Learning (e.g., co-training) is an example of a semi-supervised algorithms. Reinforcement algorithms work by providing rewards or penalties to the result of the algorithm’s suggested actions. The algorithms can learn something given by an external feedback from the environment or a thinking entity, in a continuously trial-and-error way. A commonly used reinforcement learning algorithm is Q-Learning. (Portugal et al., 2018; Sutton & Barto, 1998). 2.2.2 Auto Machine Learning (AutoML) One of the key tasks of ML is identifying a model to use for a particular dataset, the attributes to consider and defining the right choice of its hyperparameters (Feurer, Springenberg, et al., 2015). When performed 42
2.3. BUSINESS ANALYTICS APPLIED TO THE CHEMICAL INDUSTRY Table 11: Overview of the Practical Articles that used Business Analytics in the Chemical Domain Reference Func.13Data14 Goal Impact Modeling15 Montavon et al. (2013) IP L ANN model that simultaneously predicts multiple electronic groundand excited-state properties. 5 ANN Morellos et al. (2016) IP L Used Regression method to predict the soil total nitrogen, organic carbon and moisture. 4 C, LSSVM, PCR, PLSR Coley et al. (2018) IP L Used Neural Networks to improve the synthesis planning. 5 NN Häse et al. (2018) IP L Implemented an Optimization framework for self-driving laboratories. 4 SOOA Wahab et al. (2020) ID H Artificial Neural Networks to predict energy consumption at the laboratories. 5 ANN M. Zhong et al. (2020) IP L Integrated Machine Learning Algorithms in a framework to accelerate the discovery of chemical compounds. 4 ND 13Hardware Connection (HC), Information Discovery (ID), Intelligent Production (IP), Predictive Maintenance (PdM) 14Laboratory (L) 15 Artificial Neural Networks (ANN), Cubist (C), Least Squares Support Vector Machines (LS-SVM), Neural Networks (NN), Not Disclosed (ND), Principal Component Regression (PCR), Partial Least Squares Regression (PLSR), Single-objective optimization algorithms (SOOA) 49
Chapter 3 Methods, Experiments and Results This chapter presents the main methods, experiments and results obtained during this PhD work. The first section presents the framework that was developed and used as a guide to develop our IDSS. The remaining sections introduce the published articles, which are presented following a chronological order (also the same order assumed by the PhD project execution): • Section 3.2 presents the development of a two-stage ML model to predict the sample arrival at the AL. The associated work was published in the 16th International Conference on Artificial Intelligence Applications and Innovations (A. J. Silva et al., 2020). • In Section 3.3 the focus is on the prediction of material consumption at the AL. This work was published in the 17th International Conference on Artificial Intelligence Applications and Innovations (A. J. Silva & Cortez, 2021) • Finally, Section 3.4 presents the instruments allocation module, along with the development of the proposed IDSS that contains several data analytics modules integrated in dashboards. This work was submitted to a scientific conference. 3.1 Adopted Framework As mentioned earlier, the work at the AL of the analyzed Chemical company is mainly based on the use of physical documentation. Moreover, there are several IT applications and databases that work as silos, with few or none data integration. In this PhD project, we present an IDSS architecture that uses ML algorithms and other data analytics, where the objective is to improve the functioning of the AL as well as the existing workflow between Laboratories, Warehouse and Manufacturing. This IDSS is intended to give a unified view of this workflow, and also help bring new insights to support the work executed at the AL. To achieve the above goals, we assumed an integrated framework that is illustrated in Figure 9. We take this framework as an instantiation of the DSRM-IS methodology that was followed in this PhD thesis. In addition, we also use the CRISP-DM methodology for the development of the ML models (as shown 50
3.2. PREDICT SAMPLE ARRIVAL IN THE LABORATORIES in Sections 3.2 and 3.3). The first two components of the adopted framework have a parallel with the first two stages of the CRISP-DM methodology and were essential for the development of the remainder components. In effect, both Business and Data Understanding components were executed when designing the IDSS four main modules, which are: •Sample Arrival Prediction – Based on the Production and Warehouse data, the system will predict the sample arrival at the Laboratories. This applies for the RM, IPC, and FP samples. •Materials Consumptions Prediction – Using the Sample Arrival data (forecasted and historical) in the Laboratories and material requests data, the goal is to predict the Materials Consumptions in the Laboratories in order to guarantee that the Warehouse always have the quantity of the material to be requested. •Suggest Instruments Allocation – Using historical of records the instruments usage, knowledge about the samples that will be arriving at the Laboratories, the respective information regarding the product, sample type and tests to be performed, the goal of this module is to assign the best instrument for the quality analysis. •Decision Support Dashboards – This module joins all the predictions and suggestions created in the previous modules and presents them to the users using friendly Dashboards. This module also generates reports about the activities performed in the laboratory, regarding the sample arrival, material requests and instruments allocations, as well as the historical visualization of samples arrived and tests performed at the laboratory. These Dashboards are useful to find and resolve bottlenecks in the laboratory workflow. In the next sections of this chapter, the developments related with the last three components of the framework are presented. In Section 3.2 the sample arrival prediction module is detailed. Section 3.3 presents the module for predicting the consumption of materials in the AL. Finaly, Section 3.4 contains the instruments allocation module along with a presentation of the fully designed IDSS. 3.2 Predict Sample Arrival in the Laboratories In this study, we address a relevant Business Analytics need of a Chemical company, which is adopting a Industry 4.0 transformation. To ensure the quality of the products being manufactured, samples taken from the company production processes need to be tested in Laboratories. The tests assure that the products are compliant with quality standards, allowing their usage by the company clients. Under this context, predicting the arrival of production samples at the laboratory is a key issue, since it helps in the allocation of equipment and human resources. Aiming to solve this task, this study presents a novel two-stage ML prediction system, which was developed during the implementation of a CRISP-DM (Wirth & Hipp, 2000) project that included three iterations, each focusing on a distinct regression strategy. During 51
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS Business Understanding Understanding the laboratories limitation and Machine Learning Needs Data Understanding Collect the Data used by the Laboratories, Warehouse and Production Verify data quality Study Data Sample Arrival Prediction (Section 3.2) Select Data to predict sample arrival Select the Machine Learning approach Evaluate the models Predict Materials Consumptions (Section 3.3) Use the Forecasted Samples arrival with material requests data Select the Machine Learning approach Evaluate and Compare the models Suggest Instuments Allocation (Section 3.4) Use the Forecasted Samples arrival with instruments historical data Study different Business Analytics Approachs to assign the Instruments Evaluate the studied approaches Decision Support Dashboards (Section 3.4) Sample Arrival Forecasting Material Consumption Prediction Instruments Allocation Perfomance Metrics and Weekly/Monthly Reports Figure 9: Adopted PhD framework. the modeling stage of the three CRISP-DM iterations, an AutoML (Feurer, Klein, et al., 2015) procedure was adopted, allowing to compare and configure six state-of-the-art ML algorithms. 3.2.1 Materials and Methods 3.2.1.1 Business Task This Chemical company produces several products, in batches. During the production-batch execution process, a sequence of samples, called IPC, are selected for quality Laboratory inspection, in order to ensure that the production process is running as expected. In terms of the Chemical Laboratories, the IPC samples have the highest priority, because the production process can not continue without their approval. A fixed amount of IPC samples are selected from each production-batch (𝑠∈ {1, ..., 𝐼𝑃𝐶max}). The production information system registers several attributes related to the IPC sample production, including its initial production time, denoted here as IPC production time 𝑃𝑇𝑠. One by one, the IPC samples arrive at the Laboratory at time 𝐿𝑇𝑠, under irregular intervals that are difficult to be estimated in advance. The business goal is thus the non-trivial task of predicting of arrival time for each IPC sample at the Chemical Laboratories. Solving this task efficiently allows a better management of the Laboratory equipment and human resources. For instance, some IPC quality tests require a setup time, in which the analysts need to prepare in advance the Laboratory testing instruments. The business goal was 52
3.2. PREDICT SAMPLE ARRIVAL IN THE LABORATORIES Table 12: Summary of the data attributes. Input Attributes: Name Description Range day day of the week when the production-batch started {1,...,7} month month when the production-batch started {1,...,12} product product type (nominal code) 155 levels version version of the product (numeric) {1,...,108} grade product grade (nominal, related with the lab tests) 15 levels stage product stage (nominal, related with the lab tests) 1,272 levels batch batch identification of the product (nominal) 925 levels 𝑠sequence number of the sample (𝑠∈ {1, ..., 𝐼𝑃𝐶max}) {1,...,169} Output Targets: Name Description Range 𝑦1time lag arrival of two consecutive samples [0.2,5315.3] 𝑦2time lag between 𝑃𝑇𝑠and 𝐿𝑇𝑠[0.0,3270.0] addressed as a regression task, under two main target goals. In the first CRISP-DM iteration, we only used Laboratory temporal data and the target goal was defined as predict𝑦1=𝐿𝑇𝑠+1−𝐿𝑇𝑠, which corresponds to the time lag between the next IPC sample arrival (𝐿𝑇𝑠+1) and the current (already known) Laboratory sample arrival (𝐿𝑇𝑠). In the second and third CRISP-DM iterations, we explored production temporal data, predicting 𝑦2=𝐿𝑇𝑠−𝑃𝑇𝑠, where the Laboratory arrival time can be immediately estimated once the IPC sample starts its production. 3.2.1.2 Data Understanding and Preparation We used an ETL procedure to merge the relevant data from two main databases related with the production and Laboratory testing information systems, populating an integrated and business oriented data warehouse system. The ETL resulted in a raw file with 226,929 rows and 33 columns regarding all Laboratory samples that were analyzed during a three-year time period. The data warehouse was further filtered in order to contain rows related with IPC samples and with complete values in terms of the input and output attributes (Table 12), leading to a dataset with 26,611 instances. The input variables were manually selected and defined from the filtered raw file using expert domain knowledge, obtained by interacting with the chemistry experts. Due the complexity of the Chemical factory processes and information system integration issues, it was not possible to have access to a more richer set of data features (e.g., which components and machines were used to produce the samples). Thus, the resulting set of 8 inputs is rather small, which makes more challenging the prediction task. Both output targets were computed using a particular time unit, which is not disclosed here due to business privacy issues. 53
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS 3.2.1.3 Machine Learning Models In terms of computational environment, we adopted the R tool and its rminer package (Cortez, 2014) for data manipulation and ML result evaluation, while the AutoML adopts the H2O implementation (Cook, 2016). The AutoML procedure was configured to select the regression model and its hyperparameters based on the best RMSE computed using a validation set that is obtained by applying an internal 10fold cross-validation method over the training data. All computational experiments were executed on the same personal computer and each individual ML model was trained up to a maximum running time of 3,600 seconds. Once a ML model is selected, the model was retrained with all training data. As in (Ferreira et al., 2020), the AutoML was configured to include a total of 6 distinct regression algorithms: RF, Extremely Randomized Trees (XRT), Generalized Linear Model (GLM), Gradient Boosting Machine (GBM), XGBoost (XGB) and a Stacked Ensemble (SE). The RF is a popular ensemble method that combines a large number of decision trees based on bagging and random selection of input features (Hastie et al., 2009). The XRT algorithm extends the RF approach by randomly selecting the decision thresholds of the tree nodes (Geurts et al., 2006). GLM estimates regression models for exponential distributions (e.g., Gaussian, Poisson, gamma) (Hastie et al., 2009). The GBM algorithm is a based on a generalization of tree boosting, sequentially building regression trees for all data features (Hastie et al., 2009). XGB is another ensemble tree method that uses boosting to enhance the prediction results (T. Chen & Guestrin, 2016). The SE method, also known as stacked regression (Breiman, 1996), combines the predictions of different base learners by using a second-level ML algorithm. The H2O implementation (Cook, 2016) uses the following AutoML setup: RF and XRT – set with the default hyperparameters; GLM - grid search used to set one hyperparameter ( alpha , a regularization parameter); GBM and XGB – grid search used to tune nine and ten hyperparameters (e.g., number of trees, maximum depth, minimum rows); SE – all five algorithms (RF, XRT, GLM, GBM, XGB) are used as base learners and the individual predictions are weighted by using a second-level GLM learner. For the ML algorithms that require numeric inputs (e.g., GLM), the nominal inputs (e.g., product, grade) are previously transformed by using the standard one-hot encoding, which assigns one boolean input per categorical level. For instance, a categorical feature with three levels ({𝑎,𝑏,𝑐}) is encoded as: 𝑎=(1,0,0), 𝑏=(0,1,0) and 𝑐=(0,0,1). A total of three CRISP-DM iterations were executed, aiming to improve the regression results and the potential value of the ML models. The first CRISP-DM iteration targeted the 𝑦1output, while the second and third CRISP-DM iterations approached 𝑦2, under two variants. The 𝑦1target is assumes that at least one IPC sample from the production-batch as arrived at the Laboratory. The trained ML model can be used each time new sample arrives, allowing to estimate when the next sample will be delivered (b𝑦1). A different perspective is adopted by the 𝑦2target, since the fitted ML model can be applied to predict the Laboratory sample arrival once an IPC sample production has started. The model employed in the second CRISP-DM iteration uses a simple regression with a single ML model (b𝑦2). During the evaluation stage of the second CRISP-DM iteration, we identified that there were some high prediction errors, in particular when predicting the arrival times for the first sample of the production-batch (𝑠=1). In order to check 54
3.2. PREDICT SAMPLE ARRIVAL IN THE LABORATORIES examples α first IP sample? data Yes s=1 predict s=1 training examples ML fit with No with 8 inputs feed the ML with 7 inputs feed the ML β 2 y ML fit with s>1 training Figure 10: Schematic of the proposed two-stage ML prediction model (b𝑦2𝛼𝛽 ). if we could improve these results, a third CRISP-DM iteration was executed, in which we specialize two distinct ML models (𝛼and 𝛽). The first ML model (𝛼) is trained using only the first product-batch sample examples (𝑠=1) and thus the fitted model includes only seven input attributes ({day, month, product, version, grade, stage,batch}). The second model (𝛽) is only activated when producing the other productbatch IPC samples (𝑠>1). Similarly to the second CRISP-DM iteration model, this ML model is trained with all eight inputs (including 𝑠, the sample sequence number). The proposed two-stage model (b𝑦2𝛼𝛽 ) is shown in Figure 10. 3.2.1.4 Evaluation The collected data was divided into three main sets, by using a chronological order. The last 20 weeks of data (total of 5,110 examples) was kept out of the initial ML experiments. The goal is apply this additional unseen data in a more realistic evaluation, provided by a RW validation (Tashman, 2000) that is executed for the best ML regression approach. The remaining and oldest 21,501 examples (not used as test set by the RW) were further divided into training and test sets (holdout split) (Schorfheide & Wolpin, 2012). The time ordered Holdout Split (HS) was used to compare the three distinct main regression approaches (from the CRISP-DM iterations). The training data included the oldest 15,050 examples (around 70%). As for the HS test set, it included 6,451 instances. Regarding the RW, it was set using a fixed training window with six months of data and a weekly testing of the ML models, in a total of 20 iterations. In the first iteration, at the first Sunday, the ML was trained with the last six months of historical data. Then, the model was used to perform sample arrival predictions for the incoming week (fixed test size of seven days). In the second iteration, executed at the second Sunday, the training window was updated by discarding one week of the oldest data and adding 55
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS the previous week examples, allowing to update (retrain) the ML model, which then predicted the next week sample arrival times, and so on. In this work, we adopt two popular regression error measures: RMSE and MAE. We also use the Acc@𝑇 metric, which is more easily understood by the business analysts, since it measures the percentage of examples accurately predicted when assuming an absolute error tolerance of 𝑇. A quality regression model should provide low RMSE and MAE values and also a high accuracy for a small 𝑇value. The Acc@𝑇concept allows to compare the predictive performance of different regression modes in a single graph, as proposed in (Bi & Bennett, 2003) with the REC curves, which plot in the 𝑦-axis the Acc@𝑇for different 𝑇values (𝑥-axis). The overall quality (for distinct 𝑇values) can be measured by computing the AREC curve when assuming a maximum tolerance of 𝑇max (in %). 3.2.2 Results Table 13 presents the test data errors, in terms of the RMSE error measure, for the HS evaluation and when comparing the two 𝑦2prediction strategies: b𝑦2, executed during the second CRISP-DM iteration; and b𝑦2𝛼𝛽 , explored in the third CRISP-DM iteration. The RMSE values confirm that for both prediction strategies, it is more difficult to predict the arrival of the first IPC sample (𝑠=1) than the arrival of the remaining samples (𝑠>1). It is interesting to notice that by specializing a learning model for each of these IPC sample types, as executed in the third CRISP-DM iteration (b𝑦2𝛼𝛽 ), a substantial error reduction is achieved for both sample types (𝑠=1and 𝑠>1). Table 13: Test data holdout results for 𝑠=1and 𝑠>1IPC sample arrival (best values in bold). RMSE Approach 𝑠=1𝑠>1 b𝑦2209.9 188.9 b𝑦2𝛼𝛽 124.8 41.3 The full comparison of the aggregated HS results, assuming all IPC samples, is shown in Table 14, which contains: the evaluation method used (Eval.); the best model selected using the AutoML procedure (Model); and several predictive performance measures. The AREC was computed by using a maximum tolerance of 𝑇max=16 time units. All performance measures confirm that the best predictive model was achieved by b𝑦2𝛼𝛽, while b𝑦1obtained better results than b𝑦2. When compared with b𝑦1,b𝑦2𝛼𝛽 achieved a substantial predictive improvement: RMSE – reduction of 46.8 points; MAE – difference of 14.1 points; and AREC – increase of 10 percentage points. As for the ML algorithms, the AutoML selected GBM and SE as the best performing models when using the 10-fold internal cross-validation (applied over training data). The b𝑦2𝛼𝛽 uses GBM for predicting the arrival times of the 𝑠=1samples and SE for the other ones. Figure 11 complements the HS results by showing the respective REC curves for the three main regression approaches. The plot confirms that for most of the low tolerance range (𝑥-axis), b𝑦2𝛼𝛽 provides 56
3.2. PREDICT SAMPLE ARRIVAL IN THE LABORATORIES a higher classification accuracy, resulting in an overall higher AREC. Indeed, the proposed two-stage ML model can predict correctly 37%, 59% and 70% of the samples for low tolerance values of 𝑇=1,𝑇=2 and 𝑇=4, a value that increases to 85% when the tolerance is increased to 𝑇=16 time units. Table 14: Test data results (best HO values in bold). Acc@𝑇 Approach Eval. Model RMSE MAE AREC 𝑇=1 𝑇=2 𝑇=4𝑇=8𝑇=16 b𝑦1GBM 98.0 27.0 61% 28% 45% 56% 66% 76% b𝑦2HO SE 190.3 112.1 6% 1% 1% 3% 5% 12% b𝑦2𝛼𝛽 𝛼:GBM;𝛽:SE 51.2 12.9 71%37%59%70%77%84% b𝑦2𝛼𝛽 RW 𝛼:GBM;𝛽:SE 37.5 11.4 71% 38% 56% 69% 76% 85% 0 5 10 15 0 20 40 60 80 100 Tolerance (time units) Accuracy (%) y_1 (AREC=61%) y_2 (AREC=6%) y_2alphabeta (AREC=71%) Figure 11: Holdout REC curves for the three regression approaches. To estimate how the selected model (b𝑦2𝛼𝛽 ) would behave in a real environment setting, we tested it under a RW evaluation. Figure 12 presents the scheme of a RW. The results for all 20 week iterations are 57
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS Figure 12: Schematic of the Rolling Window (RW) evaluation. shown in terms of the last row of Table 14 and show consistency when compared with the HS evaluation. In effect, the same glsarec value is achieved (71%), while the RMSE and MAE values are slightly lower (RMSE of 37.5 and MAE of 11.4). This is an interesting result, since the RW evaluation used more recent test data, not seen when comparing the HS results. The obtained results were presented to the business domain experts, which considered them very positive, encouraging the incorporation of the two-stage prediction model into a friendly dashboard that included several business indicators to support the Laboratory management decisions. To facilitate the visualization, the dashboard was designed to provide different granularity levels (hourly, daily or monthly) for the sample arrival prediction. For demonstrative purposes, Figure 13 plots the real and predicted values when assuming a daily aggregation of the IPC sample arrival for a particular Chemical Laboratory and for the entire RW testing time period. Due to business privacy issues, the scale of the 𝑦-axis is omitted from the graph. Figure 13 shows that the predictions are very close to the real values, denoting a high quality fit of the prediction model. 3.3 Predict Material Consumption in the Laboratories 3.3.1 Introduction During the production process, selected samples are sent to be tested at the AL, which is responsible for assuring that the products are compliant with quality standards. The analysis of a sample at the AL requires diverse instrumental analyses, each consuming one or more materials (e.g., Acetone, Dichloromethane, Ethanol, Methanol). Under this context, predicting the amount of materials needed for the quality tests is crucial to support a AL stock management, preventing quality inspection delays which would prejudice production. In section Section 3.2, we have adopted a ML approach to successfully predict the arrival times of samples at the AL. By using this predictive approach, the Chemical organization can now perform weekly plans of the expected instrumental AL usage. Under this context, and having in account that 58
3.4. AN IDSS FOR ANALYTICAL LABORATORIES WITHIN THE INDUSTRY 4.0 CONTEXT a strong potential to improve the stock management of these materials. 3.4 An IDSS for Analytical Laboratories within the Industry 4.0 context 3.4.1 Introduction In this section, we propose an IDSS that is based on Descriptive, Predictive and Prescriptive Analytics, aiming to assist the managerial decisions of AL from a Chemical Industry that is being transformed through the Industry 4.0 concept. In previous works, we have proposed ML solutions to assist some partial AL tasks: predict the arrival time of IPC samples at the quality testing laboratories (Section 3.2); and estimate the AL materials consumption based on weekly plans of AL sample analyses (Section 3.3). In this section, we present the full IDSS that integrates both Predictive analytics, supporting the allocation of AL instruments (Prescriptive Analytics). The IDSS is also complemented with Descriptive Analytics executed over AL historical records, allowing the AL managers to better identify similarities among instruments. Prior to the Industry 4.0 transformation, the relevant digital records were spread in distinct databases, located in different departments (production and the AL), making the AL manager decisions more difficult. The proposed IDSS integrates all relevant data records into a single data repository, while also providing the business analytics results in terms of an interactive visual tool, based on dashboards. A IDSS prototype was deployed in the chemical company and then evaluated by the AL managers by using the TAM 3 (Venkatesh & Bala, 2008) and open interviews. 3.4.2 Materials and Methods 3.4.2.1 Problem Formulation As stated earlier, the compay is from the chemical sector and it includes three main areas: Warehouse, Production and AL. The Warehouse is where the raw materials are received. It is also the destination of the products produced before being shipped to the customers. The Production area is where the chemical products are manufactured. Finally, the AL are responsible for testing all products and raw materials, checking if they meet the required quality standards. Before adopting an Industry 4.0 transformation, the entire communication process between these three areas was mainly manual and there was no real-time monitoring of the industrial processes, often leading to delays in the preparation of production materials or in the analyzes performed by the AL. These delays strongly affected deadlines for production plans. Concerning the AL, these involve human analysts, instruments and several types of samples, namely RM, IPC and FP, that need to be analyzed, i.e., allocated into one or more analytical instruments. In 65
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS particular, IPC samples are are a priority because if they are not analyzed in a timely manner, the production process may stop. Each instrument allocation requires time and manual effort, to prepare and conduct the analysis and then collect the obtained results. There is an IS that records all quality test data, but such IT is mostly focused on the testing measurements and not on the AL processes. Thus, the management of the AL (e.g., human resource and instrument allocation planning, sample prioritization, prior preparation of instruments), assumes a strong manual effort, which is difficult due to the lack of a real-time data communication with the Warehouse and Production areas. 3.4.2.2 Proposed IDSS To solve the previous mentioned AL management issues, and benefiting from the Industry 4.0 transformation performed at the company, we propose an IDSS that incorporates Descriptive, Predictive and Prescriptive Analytics. The proposed IDSS architecture is depicted in Figure 16. It includes two main layers. The Big Data layer is responsible for extracting and processing data from the different databases used in the organization. Indeed, the IDSS consumes the data from the different areas and applications from the organization (e.g., Warehouse, Production, AL), resulting as the ground truth data repository for the AL. The processed data is then fed into the Data Analytics layer, which incorporates Descriptive, Predictive and Prescriptive Analytics for AL management. Dashboards Databases Big Data Layer Data Analytics Layer UC2 UC1 UC3 Intelligent Decision Support System UC4 Figure 16: Proposed Architecture. The developed tool includes two predictive models that were previously studied. Both models are based on an AutoML procedure but fed with different input attributes and training data. The proposed IDSS includes an extension of the first predictive model, termed here Use Case (UC) 1 (UC1), successfully tested for estimating the arrival of IPC samples at the ALs (Section 3.2). In this proposed IDSS, the model is adapted to perform predictions for all types of AL samples (the studied IPC and also the RM and FP). It should be noted that the predictions for the RM and FP samples ran only in one part of the hybrid 66
3.4. AN IDSS FOR ANALYTICAL LABORATORIES WITHIN THE INDUSTRY 4.0 CONTEXT model, as there is no identification of the first sample of each batch. Since the results are still very much in an embryonic state, these were not considered in the sample arrival prediction study. However, the results are attached to this thesis in Annex 1. It should be noted that each sample arrived at the AL is associated with a fixed set of quality tests to be executed. The IDSS also integrates a second predictive model (UC2) that estimates the weekly consumption of AL materials (Section 3.3). This second predictive model requires, as input, a weekly plan of quality tests to be performed, which is built in advance by adopting the UC1 predictive model. The IDSS also includes Prescriptive Analytics (UC3), which is based on sample arrival estimates (UC1) and historical records regarding previous instrument allocations, allowing to provide suggestions of future instrument allocation. Finally, the IDSS also includes Descriptive Analytics set in terms of historical associations of instruments to quality tests (UC4), allowing to identify instrument similarities. All analytics are incorporated into friendly user dashboards. Regarding the UC3, to issue recommendations of AL instruments allocation, we use a statistical approach that considers the UC1 predictions (tests to be executed) and that are matched with historical records of instrument allocation. For each required test, we assume as the “best” analytical instrument, the one currently available that has been mostly used for executing such test. An instrument is considered available if the its scheduled weekly allocation is lower than 70% (a value that was defined by the AL experts). Once an instrument is allocated, the IDSS is refreshed, with the allocation records being updated. Finally, the UC4 is based on an I × T matrix computed using historical records and that measures the total number of tests (𝑡∈𝑇) executed by an instrument (𝑖∈𝐼). Then, the known Pearson correlation is used to compute the association between two rows of the matrix (i.e, two instruments). In our dashboards, the correlation matrix (Ferré, 2009) is shown as a colored heatmap, where more similar instruments are signaled by a stronger red color. 3.4.2.3 Evaluation The proposed IDSS was developed by a research team that included both AI and Chemical company experts but not the direct AL managers. Thus, to properly evaluate the IDSS, we adopted the TAM 3 (Venkatesh & Bala, 2008), allowing to define a questionnaire that contains 10 questions and that was answered by the AL managers after experimenting the proposed tool. The questionnaire assumes the following TAM 3 constructs: Perceived Usefulness (PU), Perceived Ease of Use (PEOU), Perception of External Control (PEC), Job Relevance (REL), Output Quality (OUT), and Behavioral Intention (BI). Each question included a 5-point likert scale option for each answer, ranging from 1 (extremely disagree) to 5 (extremely agree). These questionnaires were complemented by a direct feedback from the AL managers, obtained by using open interviews in which the manager freely provided their opinions about the proposed IDSS. Furthermore, we also map the capabilities of the proposed IDSS tool, which are compared with the currently available AL informational processes (denoted as “As-Is”) (Darwish, 2011). 67
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS 3.4.3 Results 3.4.3.1 Developed IDSS Prototype The designed IDSS was written using the R language, with the ML solutions being developed using specific R R Development Core Team, 2008 packages, namely rminer Cortez, 2014, H2O AutoML Aiello et al., 2016, forecast R. Hyndman et al., 2020; R. J. Hyndman and Khandakar, 2008 and shiny (Chang et al., 2021). The IDSS was fed with real-world data from the analyzed chemical company, collected from January 2016 to May 2019 and that results from a merge of the different databases adopted by the organization. The user interface was developed using shiny and it includes three main dashboards to present the Descriptive (UC4), Predictive (UC1 and UC2) and Prescriptive (UC3) Analytics. The first dashboard presents: the expected arrival of samples and quality tests to be carried out in the current week (UC1); the expected raw material consumption (UC2); the history of quality analyzes carried out in the previous week; and an overview of the historical arrival of samples to the laboratory in the last year. The second dashboard shows the current allocation of AL instruments and suggestions on the best instrument to be used for each planned test (UC3). Finally, the last dashboard contains the correlation heatmaps based on the I × T association matrix (UC4). The first dashboard is presented in Figure 17 and it contains three components. The first one is the top bar that shows warnings about issues that could occur during the current week. This includes information about how many instruments have an expected occupation above 50%, the number of analyzes without any instrument usage history, as well as the progress of test analyzes for the current day (in Figure 17, this value is set at 0%). The second middle component includes three tables, presenting: the daily sample arrival (UC1) predictions (left table); how many analyzes are planned to be carried out on the current day (middle table); and the predicted weekly AL material consumption (UC2, right table). The third bottom component has two graphs. The first plot (bottom left) shows the number of samples that arrived at the laboratories every week by type (IPC, RM, FP), while the second graph (bottom right) displays the number of analyzes performed per week by sample type. The selection of the IDSS top menu tab allows the access to the second dashboard (Figure 18). The top left component “Analysis to be performed in this week” allows to select a quality test, refreshing the middle barplot graphs that show the instruments that are used for that specific test and sample (left) or just for that specific test (without sample specification, right plot). At the same time, the table on the top right presents the UC3 results as the suggested instrument to be assigned to that specific test analysis, along with the load work for the same instruments for that week. Finally, the bottom left table contains the information about the tests that have no historical records of instrument usage. The last and third dashboard is presented in two figures and it is related with the UC4 Descriptive Analytics. Figure 19 displays the correlation tables for a given instrument divided by two groups of instrument machines: HPLC (left table) and GC (right table). The top buttons (“Chosse HPLC/GC”) allows the user to select one instrument from the displayed list. Once the instrument is selected, a table is displayed, 68
3.4. AN IDSS FOR ANALYTICAL LABORATORIES WITHIN THE INDUSTRY 4.0 CONTEXT Figure 17: Example of the first IDSS dashboard. Figure 18: Example of the second dashboard. sorting in a descending order the correlation values of most similar instruments. The third column on the tables shows the most used test analysis for each instrument. The bottom part of the third dashboard is presented in Figure 20, which shows the instrument correlation heatmaps for each group of instruments. The heatmap provides easy visualization of the most correlated HPLC and GC instruments. 69
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS Figure 19: Example of the third dashboard (instruments correlation). Figure 20: Example of the third dashboard (instruments heatmap). 3.4.3.2 Evaluation The designed TAM 3 questionnaire is shown in Table 16. The obtained results are presented in Table 17, where each value corresponds to the average of two laboratory managers. We note that these managers correspond to IT AL staff from the analyzed chemical company and that were not directly involved in the presented research. The average responses are between 3.5 (70%) and 4 (80%), which means that laboratory managers had a positive acceptance of our IDSS. The most positive answers were related with the Perceived Usefulness (PU1 and PU2), Job Relevance (REL2) and Behavioral Intention (BI1). After obtaining the questionnaire responses, we have performed individual interviews, where the AL managers provided more specific feedback about the proposed IDSS. Regarding the first IDSS dashboard, both managers agreed that the information provided was simple and objective, being valuable to help the analysts to prepare the materials and the laboratory before the sample arrival. Turning to the second dashboard, related with the instruments load, they found it interesting but signaled the lack of information about new 70
3.5. SUMMARY Table 16: The adopted TAM 3 questionnaire. Construct Items Question Perceived Usefulness (PU) PU1 Using the Dashboards improves my performance in my job. PU2 The Dashboards are (potentially) useful in my job. Perceived Ease of Use (PEOU) PEOU1 I find the Dashboard interface to be easy to use. PEOU2 It’s easy to get the information that I want from the Dashboards. Perceptions of External Control (PEC) PEC1 I have the knowledge to use the Dashboards. Job Relevance (REL) REL1 In my job, the usage of the Dashboards is important. REL2 The use of the Dashboards is pertinent to my various job-related tasks. Output Quality (OUT) OUT1 The quality of the output I get from the Dashboards is high. OUT2 I have no difficulty telling others about the results of using the Dashboards. Behavioral Intention (BI) BI1 Assuming I had access to the Dashboard, I intend to use it. Table 17: The TAM 3 questionnaire results (average of two responses). PU1 PU2 PEOU1 PEOU2 PEC1 REL1 REL2 OUT1 OUT2 BI1 4 4 3.5 3.5 3.5 3.5 4 3.5 3.5 4 instruments and analyses. As for the third dashboard, it was considered helpful, particularly the correlation heatmap, which can be useful to identify new groups of instruments. However, such identification needs to be complemented by human domain knowledge, since there are instruments within the same group that can have different capabilities (e.g., refractive-index or infra-red). The AL managers also considered the dashboard useful to check if there a overlap between groups of instruments and if new groups of instruments could be defined. Overall, the AL managers concluded that the proposed IDSS (including its three dashboards), is valuable for planning the analyzes to be carried out on the samples, to improve the instrument allocation and to know how many analyzes will be carried out. Table 18 summarizes the main features introduced by the proposed IDSS, which substantially enhance the capabilities currently available at the AL (As-Is). 3.5 Summary In the first published work, the sample arrival prediction, presented in Section 3.2, we adressed the nontrivial task of predicting the arrival of IPC samples at Chemical Laboratories for quality testing. To solve this task, we implemented the CRISP-DM methodology under three iterations, each focusing on a different 71
CHAPTER 3. METHODS, EXPERIMENTS AND RESULTS Table 18: Comparison between the current AL (As-Is) and proposed IDSS informational processes. Capabilities As-Is IDSS Historical overview of samples arrived X X Historical overview of analysis performed X X Sample arrival prediction X(UC1) Weekly estimates of materials consumption X(UC2) Expected instruments load X(UC3) Suggested allocation of instruments X(UC3) Information of analysis without instruments X X Visualization of instrument similarities X(UC4) regression approach. During the data understanding and preparation CRISP-DM stages, we collected recent data from a chemical company, resulting in 26,611 sample arrival examples related with a three-year time period. As for the modeling stage of CRISP-DM, we employed an AutoML procedure, to automatically select and configure the best model when exploring six state-of-the-art ML algorithms. Several experiments were held. Using a time ordered HS, we compared the three main regression approaches: b𝑦1predict the time lag between the arrival of two consecutive samples (𝑦1), executed in the first CRISP-DM iteration; b𝑦2 - predict the time lag between starting the production of the sample and its arrival to the laboratory (𝑦2), explored in the second CRISP-DM iteration; and b𝑦2𝛼𝛽 - a two-stage ML model to predict 𝑦2, developed in the third CRISP-DM iteration. For all predictive performance measures, the best results were achieved at the two-stage ML model, which obtained interesting results (e.g., it can accurately predict 70% of the examples under a tolerance of 𝑇=4time units). The selected two-stage ML model (b𝑦2𝛼𝛽 ) was further evaluated using a realistic RW procedure, which considered 20 weeks of unseen data. A similar predictive performance was achieved, when compared with the HS results, showing that the proposed two-stage ML model is robust for the analyzed chemical company. The second work, detailed in Section 3.3, addresses a relevant business goal of a chemical company that is being transformed under the Industry 4.0. In particular, a ML approach was conducted, aiming to predict the needs of materials (e.g., Acetone, Ethanol) used in their AL. The ML project was conducted using the CRISP-DM methodology. At the data understanding CRISP-DM stage, we collected 177 weeks of data, from January 2016 to May 2019, involving a total of 30 quality tests and up to 26 consumed AL materials. It should be noted that the chemical company is currently capable of producing weekly quality test usage plans with a good accuracy. Thus, the regression goal is to model AL material consumption as a function of the conducted quality tests. Using the collected data, we have developed large set of regression models (total of 𝑀=26 models), which were analyzed in terms of two major sets of material selections: top 10 most consumed materials (𝑀=10) and all materials (𝑀=26). To reduce the ML analyst effort, we have employed an AutoML procedure during the CRISP-DM modeling stage, which allows to automatically select the best among six different regression algorithms. A total of three CRISP-DM iterations were executed, each exploring a different FS method. For comparison purposes, we also considered two time 72
3.5. SUMMARY series forecasting methods: ARIMA and a LSTM NN. Several computational experiments were executed, by considering a realistic RW procedure that simulated 30 training and testing iterations through time. The best overall results were achieved by the AutoML FS2 method (corresponding to the third CRISP-DM iteration), which obtained a total quantity NMAE of 6.1% (top 10 selection) and 2.6% (all materials). The predictive results were shown to the AL managers, which provided a positive feedback. Finally, in Section 3.4 we present an IDSS that was developed for the AL of a chemical company that is being transformed under the Industry 4.0 concept. The proposed IDSS includes two main layers: Big Data – responsible for extracting and processing data from different data sources, leading to a single and updated AL data repository; and Data Analytics – which includes Descriptive, Predictive and Prescriptive Analytics that aim to enhance the managerial decisions performed by AL managers. Using recent data from a real-world chemical company, in Sections 3.2 and 3.3 we have proposed two Predictive Analytics (IPC sample arrival prediction – UC1 and weekly AL materials consumption – UC2). The Data Analytics layer includes these analytics, extending the arrival prediction capabilities to all AL sample types (e.g., RM and FP). Moreover, it includes a novel Prescriptive method (UC3) for suggesting instrument allocations for quality tests based on historical records and the sample arrival predictions (UC1). Finally, it includes Descriptive Analytics regarding laboratory instrument similarities (UC4). A IDSS prototype was developed, which integrated all proposed analytics in three main interactive dashboards and used data collected from January 2016 to May 2019. The prototype was evaluated by two AL managers that were not directly involved in the IDSS design by adopting TAM 3 questionnaires and open interviews. Overall, a very positive feedback was obtained. In particular, the proposed IDSS was considered valuable to better prepare and assign instruments to samples, as well as to better estimate the ammount of quality tests that will be carried out. 73
Chapter 4 Conclusions This chapter presents the conclusions of this doctoral thesis. Initially, a summary of the entire PhD work is presented, going through the definition of the research project, the SLR performed, as well as the summary of the work done throughout the project. The main results obtained are discussed. Finally, several future research directions are disclosed. 4.1 Overview A major transformation is currently occurring due to the concept of Industry 4.0, also referred to as the fourth industrial revolution. Advances in IT, such as smart and cheaper sensors, IoT, Big Data and Business Analytics, are resulting in more integrated cyber-physical systems that can improve the production process. Business Analytics is a modern trend, defined after the 2010s and includes various Forecasting and Optimization techniques that can be used to analyze historical data and provide useful, often actionable, insights to support management decisions. These techniques applied in light of the concept of Industry 4.0, can potentially bring new insights and improvements to the productive processes. This PhD program was inserted within a R&D project funded by a private company from the chemical sector, where the objective was to develop an intelligent system based on state-of-the-art technologies within the concept of Industry 4.0, to improve its processes and efficiency. This R&D project was divided into three different WP, with this PhD being inserted within the WP3, which aimed the design and development of an IS for AL based on a Big Data Warehouse system to collect and process all data. This PhD is specifically focused in the “intelligence” part of the project, where the objective of the project would be the integration of an intelligent system that uses Business Analytics techniques that is able to analyze the historical data of the Laboratory, within the context of Industry 4.0, with the aim of extracting knowledge to improve the management of the Laboratory. An initial SLR was executed, allowing to verify that there are currently no DSS in AL, the research gap that was carried out in this work. Thus, the main objective of this thesis was to create an IDSS that would use Business Analytics techniques (Descriptive, Predictive and Prescriptive) to improve the Laboratory 74
BIBLIOGRAPHY Charest, M., Finn, R., & Dubay, R. (2018). Integration of artificial intelligence in an injection molding process for on-line process parameter adjustment. 2018 Annual IEEE International Systems Conference, SysCon 2018, Vancouver, BC, Canada, April 23-26, 2018 , 1–6. https://doi.org/10.11 09/SYSCON.2018.8369500 (cit. on pp. 30, 35, 48) Chen, H., Li, L., & Chen, Y. (2020). Explore success factors that impact artificial intelligence adoption on telecom industry in china. Journal of Management Analytics , 8 (1), 36–68. https://doi.org/10.1 080/23270012.2020.1852895 (cit. on p. 13) Chen, T., & Guestrin, C. (2016). Xgboost: A scalable tree boosting system. In B. Krishnapuram, M. Shah, A. J. Smola, C. C. Aggarwal, D. Shen, & R. Rastogi (Eds.), Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, san francisco, ca, usa, august 13-17, 2016 (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785. (Cit. on p. 54) Chen, Y.-J., & Chien, C.-F. (2018). An empirical study of demand forecasting of non-volatile memory for smart production of semiconductor manufacturing. International Journal of Production Research , 56 (13), 4629–4643. https://doi.org/10.1080/00207543.2017.1421783 (cit. on p. 30) Chen, Y., Chen, H., Gorkhali, A., Lu, Y., Ma, Y., & Li, L. (2016). Big data analytics and big data science: A survey. Journal of Management Analytics , 3 (1), 1–42. https://doi.org/10.1080/23270012.20 16.1141332 (cit. on p. 15) Chiang, L., Lu, B., & Castillo, I. (2017). Big data analytics in chemical engineering. Annual Review of Chemical and Biomolecular Engineering , 8 , 63–85. https://doi.org/10.1146/annurev-chembioeng-0 60816-101555 (cit. on pp. 16, 17) Chi-Hsien, K., & Nagasawa, S. (2019). Applying machine learning to market analysis: Knowing your luxury consumer. Journal of Management Analytics , 6 (4), 404–419. https://doi.org/10.1080/23270 012.2019.1692254 (cit. on p. 13) Chiu, Y.-C., Cheng, F.-T., & Huang, H.-C. (2017). Developing a factory-wide intelligent predictive maintenance system based on industry 4.0. Journal of the Chinese Institute of Engineers , 40 (7), 562– 571. https://doi.org/10.1080/02533839.2017.1362357 (cit. on p. 48) Choi, W., Kim, J., Kim, S., & Kim, J. (2017). A study of reference metadata classification with deep learning. International Conference on Information and Communication Technology Convergence (ICTC) , 144–146. https://doi.org/10.1109/ICTC.2017.8190961 (cit. on pp. 28, 35) Chong, D., & Shi, H. (2015). Big data analytics: A literature review. Journal of Management Analytics , 2 (3), 175–201. https://doi.org/10.1080/23270012.2015.1082449 (cit. on pp. 13, 16) Cicconi, P., Russo, A. C., Germani, M., Prist, M., Pallotta, E., & Monteriù, A. (2017). Cyber-physical system integration for industry 4.0: Modelling and simulation of an induction heating process for aluminium-steel molds in footwear soles manufacturing. IEEE 3rd International Forum on Research and Technologies for Society and Industry (RTSI) , 1–6. https://doi.org/10.1109/RTSI.2 017.8065972 (cit. on p. 28) 81
BIBLIOGRAPHY Cisotto, S., & Herzallah, R. (2018). Performance prediction using neural network and confidence intervals: A gas turbine application. 2018 IEEE International Conference on Big Data (Big Data) , 2151–2159. https://doi.org/10.1109/BigData.2018.8621919 (cit. on pp. 30, 35) Clegg, D. (2015). Evolving data warehouse and bi architectures: The big data challenge. Business Intelligence Journal , 20 (1), 19–24 (cit. on p. 16). Coley, C. W., Green, W. H., & Jensen, K. F. (2018). Machine learning in computer-aided synthesis planning. Accounts of chemical research , 51 (5), 1281–1289 (cit. on p. 49). Commission, E. (2020). Digital Europe Programme: A proposed €7.5 billion of funding for 2021-2027. Retrieved January 18, 2021, from https://ec.europa.eu/digital-single-market/en/news/digitaleurope-programme-proposed-eu75-billion-funding-2021-2027. (Cit. on p. 41) Cook, D. (2016). Practical machine learning with h2o: Powerful, scalable techniques for deep learning and ai . O’Reilly Media. (Cit. on pp. 54, 61). Cortez, P. (2014). Modern optimization with r . Springer. (Cit. on pp. 54, 61, 68). Costa, C., & Santos, M. Y. (2017). The data scientist profile and its representativeness in the european e-competence framework and the skills framework for the information age. International Journal of Information Management , 37 (6), 726–734. https://doi.org/https://doi.org/10.1016 /j.ijinfomgt.2017.07.010 (cit. on p. 41) Costa, R., Figueiras, P., Jardim-Gonçalves, R., Ramos-Filho, J., & Lima, C. (2017). Semantic enrichment of product data supported by machine learning techniques. 2017 International Conference on Engineering, Technology and Innovation (ICE/ITMC) , 1472–1479. https://doi.org/10.1109 /ICE.2017.8280056 (cit. on p. 21) Darwish, A. (2011). Business process mapping: A guide to best practice . Writescope Publishers. (Cit. on p. 67). Davenport, T. H. (2013). Analytics 3.0. Harvard business review , 91 (12), 64–72 (cit. on p. 41). Deal, J. (2013). The ten most common data mining business mistakes. https://www.elderresearch.com/ most-common-data-science-business-mistakes. (Cit. on p. 47) de Sá, A., Casimiro, A., Machado, R., & Carmo, L. (2020). Identification of data injection attacks in networked control systems using noise impulse integration. Sensors (Switzerland) , 20 (3). https: //doi.org/10.3390/s20030792 (cit. on p. 35) Diez-Olivan, A., Del Ser, J., Galar, D., & Sierra, B. (2018). Data fusion and machine learning for industrial prognosis: Trends and perspectives towards industry 4.0. Information Fusion , 50 , 92–111. https: //doi.org/10.1016/j.inffus.2018.10.005 (cit. on p. 17) Duan, L., & Xiong, Y. (2015). Big data analytics and business analytics. Journal of Management Analytics , 2 (1), 1–21. https://doi.org/10.1080/23270012.2015.1020891 (cit. on p. 16) Durakbasa, N., Bauer, J., & Poszvek, G. (2017). Advanced metrology and intelligent quality automation for industry 4.0-based precision manufacturing systems. Solid State Phenomena , 261 , 432–439. https://doi.org/10.4028/www.scientific.net/SSP.261.432 (cit. on p. 25) 82
BIBLIOGRAPHY Dutta, R., Mueller, H., & Liang, D. (2018). An interactive architecture for industrial scale prediction: Industry 4.0 adaptation of machine learning. Annual IEEE International Systems Conference (SysCon) , 1– 5. https://doi.org/10.1109/SYSCON.2018.8369547 (cit. on p. 22) Dwaraka, R., & Arunachalam, N. (2018). Investigation on non-invasive process monitoring of die sinking edm using acoustic emission signals [46th SME North American Manufacturing Research Conference, NAMRC 46, Texas, USA]. Procedia Manufacturing , 26 , 1471–1482. https://doi.org/10 .1016/j.promfg.2018.07.094 (cit. on p. 30) ESRTC. (2009). Technology readiness levels: Handbook for space applications . European Space Research; Technology Centre (ESRTC). https://books.google.pt/books?id=zzYKngEACAAJ. (Cit. on p. 24) Essien, A., & Giannetti, C. (2020). A deep learning model for smart manufacturing using convolutional lstm neural network autoencoders. IEEE Transactions on Industrial Informatics , PP , 1–1. https: //doi.org/10.1109/TII.2020.2967556 (cit. on p. 35) European Commission. (2013). Multi-annual roadmap for the contractual PPP under Horizon 2020 . (Cit. on p. 40). Ferré, J. (2009). 3.02 - regression diagnostics. In S. D. Brown, R. Tauler, & B. Walczak (Eds.), Comprehensive chemometrics (pp. 33–89). Elsevier. https://doi.org/https://doi.org/10.1016/B978-0 44452701-1.00076-4. (Cit. on p. 67) Ferreira, L., Pilastri, A., Martins, C., Santos, P., & Cortez, P. (2020). An automated and distributed machine learning framework for telecommunications risk management. In A. P. Rocha, L. Steels, & H. J. van den Herik (Eds.), Proceedings of the 12th international conference on agents and artificial intelligence, ICAART 2020, volume 2, valletta, malta, february 22-24, 2020 (pp. 99–107). SCITEPRESS. https://doi.org/10.5220/0008952800990107. (Cit. on pp. 48, 54, 59, 62) Feurer, M., Klein, A., Eggensperger, K., Springenberg, J. T., Blum, M., & Hutter, F. (2015). Efficient and robust automated machine learning. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, & R. Garnett (Eds.), Advances in neural information processing systems 28: Annual conference on neural information processing systems 2015, december 7-12, 2015, montreal, quebec, canada (pp. 2962–2970). (Cit. on p. 52). Feurer, M., Klein, A., Eggensperger, K., Springenberg, J. T., Blum, M., & Hutter, F. (2019). Auto-sklearn: Efficient and robust automated machine learning. Automated machine learning: Methods, systems, challenges (pp. 113–134). Springer International Publishing. https://doi.org/10.1007/97 8-3-030-05318-5\_6. (Cit. on p. 43) Feurer, M., Springenberg, J. T., & Hutter, F. (2015). Initializing bayesian hyperparameter optimization via meta-learning. Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence , 1128– 1135 (cit. on p. 42). Fu, Y., Ding, J., Wang, H., & Wang, J. (2018). Two-objective stochastic flow-shop scheduling with deteriorating and learning effect in industry 4.0-based manufacturing system. Applied Soft Computing , 68 , 847–855. https://doi.org/10.1016/j.asoc.2017.12.009 (cit. on pp. 36, 37) 83
BIBLIOGRAPHY Geissbauer, R., Vedso, J., & Schrauf, S. (2016). Industry 4.0: Building the digital enterprise. Retrieved from PwC Website: https://www. pwc. com/gx/en/industries/industries-4.0/landing-page/industry4.0-building-your-digital-enterprise-april-2016. pdf (cit. on pp. 12, 13). Geurts, P., Ernst, D., & Wehenkel, L. (2006). Extremely randomized trees. Machine Learning , 63 (1), 3–42. https://doi.org/10.1007/s10994-006-6226-1 (cit. on p. 54) Gibert, K., Izquierdo, J., Sànchez-Marrè, M., Hamilton, S. H., Rodrı �guez-Roda, I., & Holmes, G. (2018). Which method to use? an assessment of data mining methods in environmental data science. Environmental modelling & software , 110 , 3–27 (cit. on p. 48). Gomes, M., Silva, F., Ferraz, F., Silva, A., Analide, C., & Novais, P. (2017). Developing an ambient intelligent-based decision support system for production and control planning. Advances in Intelligent Systems and Computing , 557 , 984–994. https://doi.org/10.1007/978-3-319-534800\_97 (cit. on p. 28) Goodfellow, I., Bengio, Y., Courville, A., & Bengio, Y. (2016). Deep learning . MIT press Cambridge. (Cit. on pp. 41, 42). Gottinger, H. W., & Weimann, P. (1992). Intelligent decision support systems. Decision Support Systems , 8 (4), 317–332. https://doi.org/10.1016/0167-9236(92)90053-R (cit. on p. 44) Guo, Z., Ngai, E., Yang, C., & Liang, X. (2015). An RFID-based intelligent decision support system architecture for production monitoring and scheduling in a distributed manufacturing environment. International Journal of Production Economics , 159 , 16–28. https://doi.org/j.ijpe.2014.09.004 (cit. on p. 14) Guyon, I., Bennett, K., Cawley, G., Escalante, H. J., Escalera, S., Tin Kam Ho, Macià, N., Ray, B., Saeed, M., Statnikov, A., & Viegas, E. (2015). Design of the 2015 chalearn automl challenge. 2015 International Joint Conference on Neural Networks (IJCNN) , 1–8. https://doi.org/10.1109/IJCNN.2 015.7280767 (cit. on p. 43) Haffner, O., Kucera, E., Kozák, Š., & Stark, E. (2017). Proposal of system for automatic weld evaluation. 21st International Conference on Process Control (PC) , 440–445. https://doi.org/10.1109 /PC.2017.7976254 (cit. on p. 29) Häse, F., Roch, L. M., & Aspuru-Guzik, A. (2018). Chimera: Enabling hierarchy based multi-objective optimization for self-driving laboratories. Chemical science , 9 (39), 7642–7655 (cit. on p. 49). Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction . Springer. (Cit. on pp. 54, 62). He, M., & He, D. (2017). Deep learning based approach for bearing fault diagnosis. IEEE Transactions on Industry Applications , 53 (3), 3057–3065. https://doi.org/10.1109/TIA.2017.2661250 (cit. on p. 29) He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., & Han, S. (2018). Amc: Automl for model compression and acceleration on mobile devices. Proceedings of the European Conference on Computer Vision (ECCV) , 784–800 (cit. on p. 43). 84
BIBLIOGRAPHY Hesser, D. F., & Markert, B. (2019). Tool wear monitoring of a retrofitted cnc milling machine using artificial neural networks. Manufacturing Letters , 19 , 1–4. https://doi.org/10.1016/j.mfglet.2018.11.0 01 (cit. on p. 32) Hyndman, R., Athanasopoulos, G., Bergmeir, C., Caceres, G., Chhay, L., O’Hara-Wild, M., Petropoulos, F., Razbash, S., Wang, E., & Yasmeen, F. (2020). forecast: Forecasting functions for time series and linear models [R package version 8.13]. https://pkg.robjhyndman.com/forecast/. (Cit. on pp. 61, 68) Hyndman, R. J., & Khandakar, Y. (2008). Automatic time series forecasting: The forecast package for R. Journal of Statistical Software , 26 (3), 1–22. https://www.jstatsoft.org/article/view/v027i03 (cit. on pp. 61, 68) Jin, H., Song, Q., & Hu, X. (2019). Auto-keras: An efficient neural architecture search system. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 1946– 1956. https://doi.org/10.1145/3292500.3330648 (cit. on p. 43) Jugulum, R. (2016). Importance of data quality for analytics. In P. Sampaio & P. Saraiva (Eds.), Quality in the 21st Century: Perspectives from ASQ Feigenbaum Medal Winners (pp. 23–31). Springer International Publishing. https://doi.org/10.1007/978-3-319-21332-3\_2. (Cit. on p. 36) Kabugo, J. C., Jämsä-Jounela, S.-L., Schiemann, R., & Binder, C. (2020). Industry 4.0 based process data analytics platform: A waste-to-energy plant case study. International Journal of Electrical Power & Energy Systems , 115 , 105508. https://doi.org/10.1016/j.ijepes.2019.105508 (cit. on pp. 35, 48) Kagermann, H., Helbig, J., Hellinger, A., & Wahlster, W. (2013). Recommendations for implementing the strategic initiative industrie 4.0: Securing the future of german manufacturing industry ; final report of the industrie 4.0 working group . Forschungsunion. (Cit. on p. 13). Kammergruber, R., Robold, S., Karliç, J., & Durner, J. (2014). The future of the laboratory information system–what are the requirements for a powerful system for a laboratory data management? Clinical Chemistry and Laboratory Medicine (CCLM) , 52 (11), 225–230 (cit. on p. 47). Karakose, M., & Yaman, O. (2020). Complex fuzzy system based predictive maintenance approach in railways. IEEE Transactions on Industrial Informatics , 16 (9), 6023–6032 (cit. on p. 35). Kaupp, L., Beez, U., Hülsmann, J., & Humm, B. G. (2019). Outlier detection in temporal spatial log data using autoencoder for industry 4.0. In J. Macintyre, L. Iliadis, I. Maglogiannis, & C. Jayne (Eds.), Engineering applications of neural networks (pp. 55–65). Springer International Publishing. https: //doi.org/10.1007/978-3-030-20257-6\_5. (Cit. on p. 26) Kenessey, Z. (1987). The primary, secondary, tertiary and quaternary sectors of the economy. Review of Income and Wealth , 33 (4), 359–385 (cit. on p. 3). Kharwar, P., Verma, R., Mandal, N., & Mondal, A. (2020). Swarm intelligence integrated approach for experimental investigation in milling of multiwall carbon nanotube/polymer nanocomposites. Archive of Mechanical Engineering , 67 (3), 353–376. https://doi.org/10.24425/ame.2020 .131698 (cit. on p. 38) 85
BIBLIOGRAPHY Khatri, V., & Samuel, B. M. (2019). Analytics for managerial work. Commun. ACM , 62 (4), 100. https: //doi.org/10.1145/3274277 (cit. on p. 13) Khayyam, H., Jazar, R., Nunna, S., Golkarnarenji, G., Badii, K., Fakhrhoseini, S., Kumar, S., & Naebe, M. (2019). Pan precursor fabrication, applications and thermal stabilization process in carbon fiber production: Experimental and mathematical modelling. Progress in Materials Science , 107 , 100575. https://doi.org/10.1016/j.pmatsci.2019.100575 (cit. on pp. 36, 37, 39) Kiangala, K., & Wang, Z. (2018). Initiating predictive maintenance for a conveyor motor in a bottling plant using industry 4.0 concepts. International Journal of Advanced Manufacturing Technology , 97 (912), 3251–3271. https://doi.org/10.1007/s00170-018-2093-8 (cit. on pp. 31, 35) Kim, S., Lee, Y., Adhi Tama, B., & Lee, S. (2020). Reliability-enhanced camera lens module classification using semi-supervised regression method. Applied Sciences , 10 , 3832. https://doi.org/10.339 0/app10113832 (cit. on p. 35) Kirchen, I., Schütz, D., Folmer, J., & Vogel-Heuser, B. (2017). Metrics for the evaluation of data quality of signal data in industrial processes. 2017 IEEE 15th International Conference on Industrial Informatics (INDIN) , 819–826. https://doi.org/10.1109/INDIN.2017.8104878 (cit. on pp. 25, 26, 39) Kitchenham, B., Brereton, O. P., Budgen, D., Turner, M., Bailey, J., & Linkman, S. (2009). Systematic literature reviews in software engineering–a systematic literature review. Information and software technology , 51 (1), 7–15 (cit. on p. 18). Klement, N., & Silva, C. (2017). A generic decision support tool to planning and assignment problems: Industrial applications and industry 4.0. In B. Sokolov, D. Ivanov, & A. Dolgui (Eds.), Scheduling in industry 4.0 and cloud manufacturing (pp. 167–192). Springer International Publishing. https: //doi.org/10.1007/978-3-030-43177-8\_9. (Cit. on p. 36) Koch, R. (2015). From business intelligence to predictive analytics. Strategic Finance , 96 , 56–57 (cit. on pp. 13, 22). Kohlert, M., & König, A. (2016). Advanced multi-sensory process data analysis and on-line evaluation by innovative human-machine-based process monitoring and control for yield optimization in polymer film industry. Technisches Messen , 83 (9), 474–483. https://doi.org/10.1515/teme-2015-0120 (cit. on p. 27) Krishnamoorthi, S., & Mathew, S. K. (2018). Business analytics and business value: A comparative case study. Information & Management , 55 (5), 643–666. https://doi.org/10.1016/j.im.2018.01.00 5 (cit. on pp. 1, 12) Krishnan, K. (2013). Data warehousing in the age of big data (1st). Morgan Kaufmann Publishers Inc. https://doi.org/10.1016/C2012-0-02737-8. (Cit. on p. 16) Kumar, A., Chinnam, R. B., & Tseng, F. (2018). An HMM and polynomial regression based approach for remaining useful life and health state estimation of cutting tools. Computers & Industrial Engineering . https://doi.org/10.1016/j.cie.2018.05.017 (cit. on p. 31) 86
BIBLIOGRAPHY Kuo, C.-J., Ting, K.-C., Chen, Y.-C., Yang, D.-L., & Chen, H.-M. (2017). Automatic machine status prediction in the era of industry 4.0: Case study of machines in a spring factory. Journal of Systems Architecture , 81 , 44–53. https://doi.org/10.1016/j.sysarc.2017.10.007 (cit. on pp. 25, 26) Kuo, H., & Faricha, A. (2016). Artificial neural network for diffraction based overlay measurement. IEEE Access , 4 , 7479–7486. https://doi.org/10.1109/ACCESS.2016.2618350 (cit. on pp. 27, 35) Langone, R., Cuzzocrea, A., & Skantzos, N. (2020). Interpretable anomaly prediction: Predicting anomalous behavior in industry 4.0 settings via regularized logistic regression tools. Data & Knowledge Engineering , 130 , 101850. https://doi.org/https://doi.org/10.1016/j.datak.2020.101850 (cit. on p. 48) Larose, D. T. (2004). Discovering knowledge in data: An introduction to data mining . Wiley-Interscience. (Cit. on p. 42). Lasi, H., Fettke, P., Kemper, H.-G., Feld, T., & Hoffmann, M. (2014). Industry 4.0. Business & Information Systems Engineering , 6 , 239–242. https://doi.org/10.1007/s12599-014-0334-4 (cit. on p. 13) Lasinkas, J. (2017). Industry 4.0: Penetrating digital technologies reshape global manufacturing sector. Retrieved June 25, 2018, from https : / / blog . euromonitor . com / 2017 / 01 / industry - 4 - 0 - penetrating-digital-technologies-reshape-global-manufacturing-sector.html. (Cit. on p. 14) Lee, J., Lapira, E., Bagheri, B., & Kao, H. (2013). Recent advances and trends in predictive manufacturing systems in big data environment. Manufacturing Letters , 1 (1), 38–41. https://doi.org/10.1016 /j.mfglet.2013.09.005 (cit. on p. 15) Lee, J., Kao, H.-A., & Yang, S. (2014). Service innovation and smart analytics for industry 4.0 and big data environment [Product Services Systems and Value Creation. Proceedings of the 6th CIRP Conference on Industrial Product-Service Systems]. Procedia CIRP , 16 , 3–8. https://doi.org/10 .1016/j.procir.2014.02.001 (cit. on p. 21) Lee, W. J., Wu, H., Yun, H., Kim, H., Jun, M., & Sutheralnd, J. (2019). Predictive maintenance of machine tool systems using artificial intelligence techniques applied to machine condition data. Procedia CIRP , 80 , 506–511. https://doi.org/10.1016/j.procir.2018.12.019 (cit. on pp. 33, 35) Lee, Y.-M., Lin, W. .-., Li, M.-H., Xiangqian, Z., & Li, J.-Y. (2016). Research into real-time analysis and exploration of influences on load rate of main shaft of machine of case companies with industry 4.0 technology. 2016 International Conference on Fuzzy Theory and Its Applications (iFuzzy) , 1–7. https://doi.org/10.1109/iFUZZY.2016.8004968 (cit. on pp. 25, 26) Leite, M., Pinto, T., & Alves, C. (2019). A real-time optimization algorithm for the integrated planning and scheduling problem towards the context of industry 4.0. FME Transactions , 47 (4), 775–781. https://doi.org/10.5937/fmet1904775L (cit. on p. 37) Lenz, J., Wuest, T., & Westkämper, E. (2018). Holistic approach to machine tool data analytics [Special Issue on Smart Manufacturing]. Journal of Manufacturing Systems , 48 , 180–191. https://doi. org/10.1016/j.jmsy.2018.03.003 (cit. on pp. 25, 26) 87
BIBLIOGRAPHY Li, H. (2016). An approach to improve flexible manufacturing systems with machine learning algorithms. IECON 2016 - 42nd Annual Conference of the IEEE Industrial Electronics Society , 54–59. https: //doi.org/10.1109/IECON.2016.7793838 (cit. on p. 36) Li, S. C., Huang, Y., Tai, B. C., & Lin, C. T. (2017). Using data mining methods to detect simulated intrusions on a modbus network. IEEE 7th International Symposium on Cloud and Service Computing (SC2) , 143–148. https://doi.org/10.1109/SC2.2017.29 (cit. on pp. 29, 35) Li, Y., Carabelli, S., Fadda, E., Manerba, D., Tadei, R., & Terzo, O. (2020). Machine learning and optimization for production rescheduling in industry 4.0. International Journal of Advanced Manufacturing Technology , 110 (9-10), 2445–2463. https://doi.org/10.1007/s00170-020-05850-5 (cit. on p. 38) Li, Z., Wang, Y., & Wang, K.-S. (2017). Intelligent predictive maintenance for fault diagnosis and prognosis in machine centers: Industry 4.0 scenario. Advances in Manufacturing , 5 (4), 377–387. https: //doi.org/10.1007/s40436-017-0203-8 (cit. on p. 29) Liang, Y., Kuo, C., & Lin, C. (2019). A hybrid memetic algorithm for simultaneously selecting features and instances in big industrial iot data for predictive maintenance. IEEE 17th International Conference on Industrial Informatics (INDIN) , 1 , 1266–1270. https://doi.org/10.1109/INDIN41052.2019 .8972199 (cit. on p. 37) Lin, C., Shu, L., Deng, D., Yeh, T., Chen, Y., & Hsieh, H. (2018). A mapreduce-based ensemble learning method with multiple classifier types and diversity for condition-based maintenance with concept drifts. IEEE Cloud Computing , 1–1. https://doi.org/10.1109/MCC.2017.455160123 (cit. on p. 30) Lin, C., & Yang, J. (2018). Cost-efficient deployment of fog computing systems at logistics centers in industry 4.0. IEEE Transactions on Industrial Informatics , 14 (10), 4603–4611. https://doi.org/1 0.1109/TII.2018.2827920 (cit. on p. 25) Lin, T., Chen, Y., Yang, D., & Chen, Y. (2016). New method for industry 4.0 machine status prediction - a case study with the machine of a spring factory. International Computer Symposium (ICS) , 322–326. https://doi.org/10.1109/ICS.2016.0071 (cit. on pp. 27, 39) Liulys, K. (2019). Machine learning application in predictive maintenance. 2019 Open Conference of Electrical, Electronic and Information Sciences (eStream) , 1–4 (cit. on pp. 33, 47). Lu, Y. (2019). Artificial intelligence: A survey on evolution, models, applications and future trends. Journal of Management Analytics , 6 (1), 1–29. https://doi.org/10.1080/23270012.2019.1570365 (cit. on p. 13) Ma, C., & Li, G. (2018). Prediction and Analysis of Tertiary Industry in Financial Center of Xinjiang under ´The Belt and Road Initiative´. 10th International Conference on Measuring Technology and Mechatronics Automation (ICMTMA) , 143–146. https://doi.org/10.1109/ICMTMA.2018.00041 (cit. on p. 22) Maggipinto, M., Terzi, M., Masiero, C., Beghi, A., & Susto, G. A. (2018). A computer vision-inspired deep learning architecture for virtual metrology modeling with 2-dimensional data. IEEE Transactions on 88
BIBLIOGRAPHY Semiconductor Manufacturing , 31 (3), 376–384. https://doi.org/10.1109/TSM.2018.2849206 (cit. on pp. 30, 35) Mahmoodpour, M., Lobov, A., Lanz, M., Mäkelä, P., & Rundas, N. (2018). Role-based visualization of industrial iot-based systems. 2018 14th IEEE/ASME International Conference on Mechatronic and Embedded Systems and Applications (MESA) , 1–8. https://doi.org/10.1109/MESA.2018 .8449183 (cit. on p. 48) Manyika, J., Chui, M., Brown, B., Bughin, J., Dobbs, R., Roxburgh, C., & Hung Byers, A. (2011). Big data: The next frontier for innovation, competition, and productivity (tech. rep.). McKinsey & Company. https://www.mckinsey.com/business-functions/digital-mckinsey/our-insights/big-data-thenext-frontier-for-innovation. (Cit. on p. 15) Martinek, P., & Krammer, O. (2019). Analysing machine learning techniques for predicting the hole-filling in pin-in-paste technology. Computers and Industrial Engineering , 136 , 187–194. https://doi.org/1 0.1016/j.cie.2019.07.033 (cit. on p. 33) Masoudinejad, M., Venkatapathy, A. K. R., Tondorf, D., Heinrich, D., Falkenberg, R., & Buschhoff, M. (2018). Machine learning based indoor localisation using environmental data in phynetlab warehouse. Smart SysTech 2018, European Conference on Smart Objects, Systems and Technologies , 1–8 (cit. on p. 21). Massaro, A., Manfredonia, I., Galiano, A., Pellicani, L., & Birardi, V. Sensing and quality monitoring facilities designed for pasta industry including traceability, image vision and predictive maintenance. In: Institute of Electrical; Electronics Engineers Inc., 2019, 68–72. https://doi.org/10.1109/ METROI4.2019.8792912 (cit. on p. 33). Massaro, A., Manfredonia, I., Galiano, A., & Xhahysa, B. (2019). Advanced process defect monitoring model and prediction improvement by artificial neural network in kitchen manufacturing industry: A case of study. IEEE International Workshop on Metrology for Industry 4.0 and IoT, MetroInd 4.0 and IoT 2019 - Proceedings , 64–67. https://doi.org/10.1109/METROI4.2019.8792872 (cit. on p. 33) Mell, P. M., & Grance, T. (2011). SP 800-145. The NIST Definition of Cloud Computing (tech. rep.). National Institute of Standards & Technology. Gaithersburg, MD, United States. (Cit. on p. 15). Michalewicz, Z., Schmidt, M., Michalewicz, M., & Chiriac, C. (2006). Adaptive business intelligence . Springer. (Cit. on p. 44). Milošević, M., Durdev, M., Lukić, D., Antić, A., & Ungureanu, N. (2020). Intelligent process planning for smart factory and smart manufacturing. Lecture Notes in Mechanical Engineering , 205–214. https://doi.org/10.1007/978-3-030-46212-3\_14 (cit. on p. 38) Miškuf, M., & Zolotová, I. (2016). Comparison between multi-class classifiers and deep learning with focus on industry 4.0. Cybernetics Informatics (K I) , 1–5. https://doi.org/10.1109/CYBERI.2016.74 38633 (cit. on pp. 28, 35) Mitchell, T. (1997). Machine learning . McGraw-Hill. (Cit. on p. 42). 89
BIBLIOGRAPHY Mohanty, S., Jagadeesh, M., & Srivatsa, H. (2013). Big data imperatives: Enterprise big data warehouse, bi implementations and analytics (1st). Apress. (Cit. on p. 16). Montavon, G., Rupp, M., Gobre, V., Vazquez-Mayagoitia, A., Hansen, K., Tkatchenko, A., Müller, K.-R., & Anatole von Lilienfeld, O. (2013). Machine learning of molecular electronic properties in chemical compound space. New Journal of Physics , 15 (9), 095003. https://doi.org/10.1088/1367-263 0/15/9/095003 (cit. on p. 49) Morellos, A., Pantazi, X.-E., Moshou, D., Alexandridis, T., Whetton, R., Tziotzios, G., Wiebensohn, J., Bill, R., & Mouazen, A. M. (2016). Machine learning based prediction of soil total nitrogen, organic carbon and moisture content by using vis-nir spectroscopy [Proximal Soil Sensing – Sensing Soil Condition and Functions]. Biosystems Engineering , 152 , 104–116. https://doi.org/https://doi.org/10.10 16/j.biosystemseng.2016.04.018 (cit. on p. 49) Moro, S., Laureano, R., & Cortez, P. (2011). Using data mining for bank direct marketing: An application of the crisp-dm methodology. Proceedings of European Simulation and Modelling ConferenceESM’2011 , 117–121 (cit. on p. 47). Mozgova, I., Yanchevskyi, I., Gerasymenko, M., & Lachmayer, R. (2018). Mobile automated diagnostics of stress state and residual life prediction for a component under intensive random dynamic loads [4th International Conference on System-Integrated Intelligence: Intelligent, Flexible and Connected Systems in Products and Production]. Procedia Manufacturing , 24 , 210–215. https: //doi.org/10.1016/j.promfg.2018.06.037 (cit. on pp. 26, 39) Muhuri, P., Shukla, A., & Abraham, A. (2019). Industry 4.0: A bibliometric analysis and detailed overview. Engineering Applications of Artificial Intelligence , 78 , 218–235. https://doi.org/10.1016/j. engappai.2018.11.007 (cit. on p. 17) Mulrennan, K., Donovan, J., Creedon, L., Rogers, I., Lyons, J. G., & McAfee, M. (2018). A soft sensor for prediction of mechanical properties of extruded pla sheet using an instrumented slit die and machine learning algorithms. Polymer Testing , 69 , 462–469. https://doi.org/10.1016/j. polymertesting.2018.06.002 (cit. on p. 30) Naskos, A., Gounaris, A., Metaxa, I., & Köchling, D. (2019). Detecting anomalous behavior towards predictive maintenance. In H. A. Proper & J. Stirna (Eds.), Advanced information systems engineering workshops (pp. 73–82). Springer International Publishing. https://doi.org/10.1007/978-3-030 -20948-3\_7. (Cit. on p. 34) Negri, E., Ardakani, H., Cattaneo, L., Singh, J., MacChi, M., & Lee, J. (2019). A digital twin-based scheduling framework including equipment health index and genetic algorithms. IFAC–PapersOnLine , 52 (10), 43–48. https://doi.org/10.1016/j.ifacol.2019.10.024 (cit. on p. 37) Neuböck, T., & Schrefl, M. (2015). Modelling knowledge about data analysis processes in manufacturing. IFAC-PapersOnLine , 48 (3), 277–282. https://doi.org/10.1016/j.ifacol.2015.06.094 (cit. on pp. 24, 26, 48) 90
BIBLIOGRAPHY Sun, I.-C., & Chen, K.-S. (2017). Development of signal transmission and reduction modules for status monitoring and prediction of machine tools. 56th Annual Conference of the Society of Instrument and Control Engineers of Japan (SICE) , 711–716. https://doi.org/10.23919/SICE.2017.81054 59 (cit. on p. 29) Susto, G. A., Schirru, A., Pampuri, S., Beghi, A., & DeNicolao, G. (2018). A hidden-gamma model-based filtering and prediction approach for monotonic health factors in manufacturing. Control Engineering Practice , 74 , 84–94. https://doi.org/10.1016/j.conengprac.2018.02.011 (cit. on p. 31) Sutton, R. S., & Barto, A. G. (1998). Introduction to reinforcement learning (1st). MIT Press. (Cit. on p. 42). Swamy, A. K., & Sarojamma, B. (2020). Bank transaction data modeling by optimized hybrid machine learning merged with arima. Journal of Management Analytics , 7 (4), 624–648. https://doi.org/1 0.1080/23270012.2020.1726217 (cit. on p. 13) Tan, Y., Goddard, S., & Pérez, L. C. (2008). A prototype architecture for cyber-physical systems. SIGBED Rev. , 5 (1), 26:1–26:2. https://doi.org/10.1145/1366283.1366309 (cit. on p. 14) Tang, D., Zheng, K., Zhang, H., Sang, Z., Zhang, Z., Xu, C., Espinosa-Oviedo, J. A., Vargas-Solar, G., & Zechinelli-Martini, J.-L. (2016). Using Autonomous Intelligence to Build a Smart Shop Floor [The 9th International Conference on Digital Enterprise Technology - Intelligent Manufacturing in the Knowledge Economy Era]. Procedia CIRP , 56 , 354–359. https://doi.org/10.1016/j.procir.201 6.10.039 (cit. on pp. 25, 26) Tashman, L. J. (2000). Out-of-sample tests of forecasting accuracy: An analysis and review [The M3Competition]. International Journal of Forecasting , 16 (4), 437–450. https://doi.org/10.1016 /S0169-2070(00)00065-0 (cit. on pp. 55, 62) Teschemacher, U., & Reinhart, G. (2017). Ant colony optimization algorithms to enable dynamic milkrun logistics [Manufacturing Systems 4.0 – Proceedings of the 50th CIRP Conference on Manufacturing Systems]. Procedia CIRP , 63 , 762–767. https://doi.org/10.1016/j.procir.2017.03.125 (cit. on p. 22) Thornton, C., Hutter, F., Hoos, H. H., & Leyton-Brown, K. (2013). Auto-weka: Combined selection and hyperparameter optimization of classification algorithms. Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , 847–855. https://doi.org/1 0.1145/2487575.2487629 (cit. on p. 43) Tieng, H., Tsai, T., Chen, C., Yang, H., Huang, J., & Cheng, F. (2018). Automatic virtual metrology and deformation fusion scheme for engine-case manufacturing. IEEE Robotics and Automation Letters , 3 (2), 934–941. https://doi.org/10.1109/LRA.2018.2792690 (cit. on p. 26) Tiwari, K., Shaik, A., & N, A. (2018). Tool wear prediction in end milling of ti-6al-4v through kalman filter based fusion of texture features and cutting forces [46th SME North American Manufacturing Research Conference, NAMRC 46, Texas, USA]. Procedia Manufacturing , 26 , 1459–1470. https: //doi.org/10.1016/j.promfg.2018.07.095 (cit. on p. 31) 97
BIBLIOGRAPHY Tjahjono, B., Esplugues, C., Ares, E., & Pelaez, G. (2017). What does industry 4.0 mean to supply chain? [Manufacturing Engineering Society International Conference 2017, MESIC 2017, 28-30 June 2017, Vigo (Pontevedra), Spain]. Procedia Manufacturing , 13 , 1175–1182. https://doi.org/1 0.1016/j.promfg.2017.09.191 (cit. on pp. 13, 14) Tran, M.-Q., Elsisi, M., Mahmoud, K., Liu, M.-K., Lehtonen, M., & Darwish, M. M. F. (2021). Experimental setup for online fault diagnosis of induction machines via promising iot and machine learning: Towards industry 4.0 empowerment. IEEE Access , 9 , 115429–115441. https://doi.org/10.110 9/ACCESS.2021.3105297 (cit. on p. 48) Trunzer, E., Weiß, I., Folmer, J., Schrüfer, C., Vogel-Heuser, B., Erben, S., Unland, S., & Vermum, C. (2017). Failure mode classification for control valves for supporting data-driven fault detection. 2017 IEEE International Conference on Industrial Engineering and Engineering Management (IEEM) , 2346– 2350. https://doi.org/10.1109/IEEM.2017.8290311 (cit. on pp. 25, 26, 39) Tsai, S., & Chang, J. J. (2018). Parametric study and design of deep learning on leveling system for smart manufacturing. IEEE International Conference on Smart Manufacturing, Industrial Logistics Engineering (SMILE) , 48–52. https://doi.org/10.1109/SMILE.2018.8353980 (cit. on p. 31) Tsourma, M., Zikos, S., Drosou, A., & Tzovaras, D. (2018). Online task distribution simulation in smart factories. 2nd International Symposium on Small-scale Intelligent Manufacturing Systems (SIMS) , 1–6. https://doi.org/10.1109/SIMS.2018.8355301 (cit. on pp. 36, 37) Uhlmann, E., Hohwieler, E., & Geisert, C. (2017). Intelligent production systems in the era of Industrie 4.0–changing mindsets and business models. Journal of Machine Engineering , 17 (cit. on pp. 16, 17). Uriarte, A. G., Ng, A. H., & Moris, M. U. (2018). Supporting the lean journey with simulation and optimization in the context of industry 4.0 [Proceedings of the 8th Swedish Production Symposium (SPS 2018)]. Procedia Manufacturing , 25 , 586–593. https://doi.org/10.1016/j.promfg.2018.06.097 (cit. on pp. 36, 37, 39) Vaishnavi, V., & Kuechler, B. (2004). Design science research in information systems. Association for Information Systems (cit. on p. 7). Vathoopan, M., Johny, M., Zoitl, A., & Knoll, A. (2018). Modular fault ascription and corrective maintenance using a digital twin [16th IFAC Symposium on Information Control Problems in Manufacturing INCOM 2018]. IFAC-PapersOnLine , 51 (11), 1041–1046. https://doi.org/10.1016/j.ifacol.2018 .08.470 (cit. on p. 26) Vazan, P., Janikova, D., Tanuska, P., Kebisek, M., & Cervenanska, Z. (2017). Using data mining methods for manufacturing process control. IFAC-PapersOnLine , 50 (1), 6178–6183. https://doi.org/10.1 016/j.ifacol.2017.08.986 (cit. on p. 29) Venkatesh, V., & Bala, H. (2008). Technology acceptance model 3 and a research agenda on interventions. Decis. Sci. , 39 (2), 273–315. https://doi.org/10.1111/j.1540-5915.2008.00192.x (cit. on pp. 8, 65, 67) 98
BIBLIOGRAPHY Ventura, F., Proto, S., Apiletti, D., Cerquitelli, T., Panicucci, S., Baralis, E., Macii, E., & Macii, A. (2019). A New Unsupervised Predictive-Model Self-Assessment Approach That SCALEs. IEEE International Congress on Big Data , 144–148. https://doi.org/10.1109/BigDataCongress.2019.00033 (cit. on p. 26) Wahab, N., mat yasin, Z., Salim, N., & Ab Aziz, N. F. (2020). Artificial neural network based technique for energy management prediction. Indonesian Journal of Electrical Engineering and Computer Science , 17 , 94. https://doi.org/10.11591/ijeecs.v17.i1.pp94-101 (cit. on p. 49) Wan, J., Tang, S., Li, D., Wang, S., Liu, C., Abbas, H., & Vasilakos, A. V. (2017). A manufacturing big data solution for active preventive maintenance. IEEE Transactions on Industrial Informatics , 13 (4), 2039–2047. https://doi.org/10.1109/TII.2017.2670505 (cit. on p. 29) Wang, Y., Tercan, H., Thiele, T., Meisen, T., Jeschke, S., & Schulz, W. (2017). Advanced data enrichment and data analysis in manufacturing industry by an example of laser drilling process. ITU Kaleidoscope: Challenges for a Data-Driven Society (ITU K) , 1–5. https://doi.org/10.23919/ITU-WT.2 017.8246990 (cit. on pp. 25, 26, 39) Wang, Y.-M., Wang, Y.-S., & Yang, Y.-F. (2010). Understanding the determinants of rfid adoption in the manufacturing industry. Technological Forecasting and Social Change , 77 (5), 803–815. https: //doi.org/10.1016/j.techfore.2010.03.006 (cit. on p. 14) Wen, Z., Xie, L., Fan, Q., & Feng, H. (2020). Long term electric load forecasting based on ts-type recurrent fuzzy neural network model. Electric Power Systems Research , 179 , 106106. https://doi.org/ https://doi.org/10.1016/j.epsr.2019.106106 (cit. on p. 48) Wirth, R., & Hipp, J. (2000). Crisp-dm: Towards a standard process model for data mining. Proceedings of the 4th international conference on the practical applications of knowledge discovery and data mining , 29–39 (cit. on pp. 47, 51, 59). Wolfe, M. (1955). The concept of economic sectors. The Quarterly Journal of Economics , 69 (3), 402–420 (cit. on p. 3). Wu, W., Zheng, Y., Chen, K., Wang, X., & Cao, N. (2018). A Visual Analytics Approach for Equipment Condition Monitoring in Smart Factories of Process Industry. IEEE Pacific Visualization Symposium (PacificVis) , 140–149. https://doi.org/10.1109/PacificVis.2018.00026 (cit. on p. 32) Xia, F., Yang, L. T., Wang, L., & Vinel, A. (2012). Internet of things. Int. J. Commun. Syst. , 25 (9), 1101– 1102. https://doi.org/10.1002/dac.2417 (cit. on p. 14) Xu, L. D., He, W., & Li, S. (2014). Internet of things in industries: A survey. IEEE Transactions on Industrial Informatics , 10 (4), 2233–2243. https://doi.org/10.1109/TII.2014.2300753 (cit. on p. 14) Xu, X., & Hua, Q. (2017). Industrial big data analysis in smart factory: Current status and research strategies. IEEE Access , 5 , 17543–17551. https://doi.org/10.1109/ACCESS.2017.2741105 (cit. on p. 17) Xu, X. (2012). From cloud computing to cloud manufacturing. Robotics and Computer-Integrated Manufacturing , 28 (1), 75–86. https://doi.org/10.1016/j.rcim.2011.07.002 (cit. on p. 15) 99
BIBLIOGRAPHY Yan, H., Wan, J., Zhang, C., Tang, S., Hua, Q., & Wang, Z. (2018). Industrial big data analytics for prediction of remaining useful life based on deep learning. IEEE Access , 6 , 17190–17197. https://doi.org/1 0.1109/ACCESS.2018.2809681 (cit. on p. 32) Yan, J., Meng, Y., Lu, L., & Li, L. (2017). Industrial big data in an industry 4.0 environment: Challenges, schemes, and applications for predictive maintenance. IEEE Access , 5 , 23484–23491. https: //doi.org/10.1109/ACCESS.2017.2765544 (cit. on p. 29) Yang, H., & Tate, M. (2009). Where are we at with Cloud Computing?: A Descriptive Literature Review. ACIS 2009 Proceedings - 20th Australasian Conference on Information Systems , 807–819 (cit. on p. 15). Yang, J., Chen, Y., Huang, W., & Li, Y. (2017). Survey on artificial intelligence for additive manufacturing. 23rd International Conference on Automation and Computing (ICAC) , 1–6. https://doi.org/10.2 3919/IConAC.2017.8082053 (cit. on p. 17) Yeh, W., Lai, C., & Tsai, J. (2019). Simplified swarm optimization for optimal deployment of fog computing system of industry 4.0 smart factory. Journal of Physics: Conference Series , 1411 (1). https:// doi.org/10.1088/1742-6596/1411/1/012005 (cit. on p. 38) Zenisek, J., Wolfartsberger, J., Sievi, C., & Affenzeller, M. (2019). Modeling sensor networks for predictive maintenance. In C. Debruyne, H. Panetto, W. Guédria, P. Bollen, I. Ciuciu, & R. Meersman (Eds.), On the Move to Meaningful Internet Systems: OTM 2018 Workshops (pp. 184–188). Springer International Publishing. https://doi.org/10.1007/978-3-030-11683-5\_20. (Cit. on p. 34) Zhang, Q., Cheng, L., & Boutaba, R. (2010). Cloud computing: State-of-the-art and research challenges. Journal of Internet Services and Applications , 1 (1), 7–18. https://doi.org/10.1007/s13174-01 0-0007-6 (cit. on p. 15) Zhang, T., Feng, Y., & Hao, B. Industrial intelligent forecast of tft-lcd based on r-svm. In: Institute of Electrical; Electronics Engineers Inc., 2019, 25–30. https://doi.org/10.1109/ICIAICT.2019.87 84848 (cit. on p. 34). Zheng, M., & Wu, K. (2017). Smart spare parts management systems in semiconductor manufacturing. Industrial Management & Data Systems , 117 (4), 754–763. https://doi.org/10.1108/IMDS-062016-0242 (cit. on pp. 25, 26, 39) Zhong, M., Tran, K., Min, Y., Wang, C., Wang, Z., Dinh, C.-T., De Luna, P., Yu, Z., Rasouli, A. S., Brodersen, P., et al. (2020). Accelerated discovery of co2 electrocatalysts using active machine learning. Nature , 581 (7807), 178–183 (cit. on p. 49). Zhong, R. Y., Xu, X., Klotz, E., & Newman, S. T. (2017). Intelligent manufacturing in the context of industry 4.0: A review. Engineering , 3 (5), 616–630. https://doi.org/10.1016/J.ENG.2017.05.015 (cit. on pp. 14–16) Zhou, H., & Yu, K. (2017). Imbalanced data classification for defective product prediction based on industrial wireless sensor network. Sixth International Conference on Future Generation Communication Technologies (FGCT) , 1–6. https://doi.org/10.1109/FGCT.2017.8103728 (cit. on p. 29) 100
BIBLIOGRAPHY Zhu, X., Ghahramani, Z., & Lafferty, J. (2003). Semi-supervised learning using gaussian fields and harmonic functions. Proceedings of the Twentieth International Conference on International Conference on Machine Learning , 912–919 (cit. on p. 42). 101
Annex I Annex 1 RM and FP first experimental results Table 19: Test data results from the first experimental test to predict the arrival of FP and RM samples. Sample Type RMSE MAE 𝑇=1 𝑇=2 𝑇=4𝑇=8𝑇=24 𝑇=48 RM 174.25 77.74 0.16% 0.58% 1.09% 6.54% 33.58% 67.35% FP 139.18 62.10 6.24% 11.01% 23.49% 44.04% 68.89% 76.33% 102