scieee AI-readable full text Open interactive document viewer

Using the IBM SPSS SW tool with wavelet transformation for CO2 prediction within IoT in Smart Home Care

Vaňuš, Jan

Abstract

Standard solutions for handling a large amount of measured data obtained from intelligent buildings are currently available as software tools in IoT platforms. These solutions optimize the operational and technical functions managing the quality of the indoor environment and factor in the real needs of residents. The paper examines the possibilities of increasing the accuracy of CO2 predictions in Smart Home Care (SHC) using the IBM SPSS software tools in the IoT to determine the occupancy times of a monitored SHC room. The processed data were compared at daily, weekly and monthly intervals for the spring and autumn periods. The Radial Basis Function (RBF) method was applied to predict CO2 levels from the measured indoor and outdoor temperatures and relative humidity. The most accurately predicted results were obtained from data processed at a daily interval. To increase the accuracy of CO2 predictions, a wavelet transform was applied to remove additive noise from the predicted signal. The prediction accuracy achieved in the selected experiments was greater than 95%.

Full text

sensors Article Using the IBM SPSS SW Tool with Wavelet Transformation for CO2Prediction within IoT in Smart Home Care Jan Vanus * , Jan Kubicek, Ojan M. Gorjani and Jiri Koziorek Department of Cybernetics and Biomedical Engineering, Faculty of Electrical Engineering and Computer Science, VSB—Technical University of Ostrava, Ostrava 70800, Czech Republic; [email protected]; (J.K.) [email protected] (O.M.G.); [email protected] (J.K.) *Correspondence: [email protected]; Tel.: +420-59-732-5856 Received: 31 January 2019; Accepted: 13 March 2019; Published: 21 March 2019   Abstract: Standard solutions for handling a large amount of measured data obtained from intelligent buildings are currently available as software tools in IoT platforms. These solutions optimize the operational and technical functions managing the quality of the indoor environment and factor in the real needs of residents. The paper examines the possibilities of increasing the accuracy of CO 2 predictions in Smart Home Care (SHC) using the IBM SPSS software tools in the IoT to determine the occupancy times of a monitored SHC room. The processed data were compared at daily, weekly and monthly intervals for the spring and autumn periods. The Radial Basis Function (RBF) method was applied to predict CO 2 levels from the measured indoor and outdoor temperatures and relative humidity. The most accurately predicted results were obtained from data processed at a daily interval. To increase the accuracy of CO 2 predictions, a wavelet transform was applied to remove additive noise from the predicted signal. The prediction accuracy achieved in the selected experiments was greater than 95%. Keywords: Smart Home Care (SHC); monitoring; prediction; trend detection; Artificial Neural Network (ANN); Radial Basis Function (RBF); Wavelet Transformation (WT); SPSS (Statistical Package for the Social Sciences) IBM; IoT (Internet of Things); Activities of Daily Living (ADL) 1. Introduction An intelligent building is one that is responsive to the requirements of occupants, organizations, and society. An intelligent building requires real-time information about its occupants so that it can continually adapt and respond [ 1 ]. Intelligent buildings respond to the needs of occupants and society, promoting the well-being of those living and working in them [ 2 ]. The researchers point out that the Adaptive House (the concept of a home which programs itself), Learning Homes and Attentive Homes must be programmed for a particular family and home and updated in line with changes in their lifestyle. The system monitors actions taken by the residents and looks for patterns in the environment which reliably predict these actions, where a neural network learns these patterns and the system then performs the learned actions automatically for improving the Quality of Life (QoL) [ 3 ]. Privacy, reliability and false alarms are the main challenges to be considered for the development of efficient systems to detect and classify the Activities of Daily Living (ADL) and Falls [ 4 ]. In order to provide a user-friendly environment for the management of the operational and technical functions along with providing support for the independent housing of senior citizens and disabled persons in buildings indicated as Smart House Care (SHC), it is necessary to make appropriate visualization of the technological process as required by the users with the possibility of the indirect monitoring of Sensors 2019,19, 1407; doi:10.3390/s19061407 www.mdpi.com/journal/sensors Sensors 2019,19, 1407 2 of 28 the seniors’ life activities based on the information obtained from the sensors used for the common management of the operational and technical functions in SHC. The properly devised visualization complements the final correct functionality of the intelligent building. Beaudin et al. described the use of computational and sensor technology in intelligent buildings with a focus on health monitoring in the households. This work comprises of a visualization for displaying health data and a proposal for improving the health and wellbeing of the users [ 5 ]. Booysen et al. explored machine-machine communication (M2M) to address the need for autonomous control of remote and distributed mobile systems in intelligent buildings [ 6 ]. Basu et al. used miniature wireless sensors in the wireless network to track and recognize the behavior of persons in the house. The visualization includes sensor data from the building to capture the duration of the activities [ 7 ]. Fleck et al. described a system based on intelligent cameras for 24-h monitoring and supervision of senior citizens. On this occasion, visualization is used to display relevant life information in this intelligent environment, which includes the evaluation of the seniors’ position and recognition of life activities [ 8 ]. The current trend for processing large volumes of measured quantities using Soft Computing Methods (SC) [ 9 , 10 ] is to use the available Big Data Analysis tools within the IoT platform [11,12]. The Internet of Things, shortly known as IoT, can be assumed as an integration layer which creates an interconnection of several physical devices, sensors, actuators, and controllers [ 13 ]. In simple words, the IoT allows objects other than computers or smartphones to use the Internet for sending and receiving information [ 14 ]. The number of connected IoT devices is rapidly growing. This growth could be the result of a very wide range of applications, ranging from basic home appliances and security systems to more sophisticated applications. For example, Xu et al. used IoT in order to construct a real-time system monitoring system for micro-environment parameters such as temperature, humidity, PM10 and PM2.5 [ 15 ]. Q. Min et al. suggested using IoT for monitoring of discrete manufacturing process based on IoT [ 16 ]. Wang et al. presented a feasible and reliable plantation monitoring system based on the Internet of Things, that combined the wireless sensor network, embedded development, GPRS communication technology, web service, and Android mobile platform [ 17 ]. Windarto et al. presented an application of IoT by implementing automation of lights and door in a room [ 18 ]. Coelho et al. used IoT to collect data from multiple heterogeneous sensors that were providing different types of information at a variety of locations in a smart home [ 19 ]. Data collection in this kind of examples usually results in big data. The Oxford dictionary defines big data as “extremely large data sets that may be analyzed computationally to reveal patterns, trends, and associations, especially relating to human behavior and interactions” [20]. One of the possibilities that come with big data processing is a predictive analysis which can provide predictions about the future or otherwise unknown events. Predictive analysis offers a wide range of applications such as social networking, healthcare, mobility, insurance, finance marketing, etc. Nyce suggested to use the predictive analysis for risk identifications and probabilities in order to provide an appropriate insurance rate, his method took advantage of marketing records, underwriting records, and claims records [ 21 ]. Predictive analysis includes many different statistical techniques ranging from data mining, predictive modeling to machine learning. Predictive modeling may be applied to many areas such as weather forecasting, business, Bayesian spam filters, advertising and marketing, fraud detection, etc. Predictive models are based on variables that are most likely to influence the outcome [ 22 ]. These variables are also known as predictors. There are many types of predictive models, such as neural networks and Decision trees. Ahmad et al. compared different methods of predictive modeling for solar thermal energy systems such as random forest, extra trees, and regression trees [23]. IBM offers a variety of services in terms of predictive analysis such as Watson analytics and SPSS [ 24 – 36 ]. AlFaris et al. reviewed the smart technologies; the interface and integration of the meters, sensors and monitoring systems with the home energy management system (HEMS) within the IoT with the outline that the smart home in practice provides the ability to the house to be net-zero energy building. Especially that it reduces the power demand and improve the energy performance by 37% better than ASHRAE standards for family villas sector [ 37 ]. Alirezaie et al. presents a Sensors 2019,19, 1407 3 of 28 framework called E-care@home, consisting of an IoT infrastructure, which provides information with an unambiguous, shared meaning across IoT devices, end-users, relatives, health and care professionals and organizations and demonstrates the proposed framework using an instantiation of a smart environment that is able to perform context recognition based on the activities and the events occurring in the home [ 38 ]. Bassoli et al. introduced a new system architecture suitable for human monitoring based on Wi-Fi connectivity, where the proposed solution lowers costs and the implementation burden by using the Internet connection that leans on standard home modem-routers, already present normally in the homes, and reducing the need for range extenders thanks to the long range of the Wi-Fi signal with energy savings of up to 91% [ 39 ]. Catherwood et al. presented an advanced Internet of Things point-of-care bio-fluid analyzer; a LoRa/Bluetooth-enabled electronic reader for biomedical strip-based diagnostics system for personalized monitoring, where practical hurdles in establishing an Internet of Medical Things network, assisting informed deployment of similar future systems are solved [40]. In this paper, the authors focused on designing a methodology for processing data obtained from sensors that measure non-electrical quantities in an SHC environment for the purpose of monitoring the presence of people in a room through the KNX (Konnex bus) and BACnet (Building Automation and Control Networks) technologies designed for SHC automation. Standard solutions using Software (SW) tools in IoT platforms are currently available and can process large amounts of measured data. These solutions monitor and optimize the quality of the indoor environment, taking into account the real needs of residents. This paper examines the possibilities of increasing the accuracy in CO 2 predictions in Smart Home Care (SHC) using IBM SPSS software tools in the IoT to determine occupancy times of a monitored SHC room. The accuracy of CO 2 predictions from the processed data was compared and evaluated at daily, weekly and monthly intervals for the spring and autumn periods. To predict CO 2 levels from the measured indoor and outdoor temperatures and relative humidity, the Radial Basis Function (RBF) method (feedforward neural network) was applied. To improve the accuracy of CO 2 predictions, a wavelet transform was applied to remove additive noise from the predicted signal. For the classification of prediction quality with Wavelet Transformation (WT) additive noise cancelation, a correlation analysis (correlation coefficient R), calculated MSE (Mean Squared Error), Mean Absolute Error (MAE), Euclidean distance (ED), City Block distance (CB) are used. 2. Description of Used Technologies in SHC The Smart two-floor wooden house (hereafter, Smart Home; floor area of 12.1 m × 8.2 m; (Figure 1) was built as a training center of the Moravian-Silesian Wood Cluster (MSWC). The wooden house (SHC) was built to a passive standard in accordance with standards ˇ CSN 75 0540-2 and ˇ CSN 730540-2 (2002) [ 41 ]. The measured values for the internal climate monitoring in the living space were evaluated by selection and measurement of CO 2 , the temperature (T) and relative humidity (rH) in room 204 with the application of air-quality sensors Siemens QPA2062 implemented in the BACnet system [ 42 ]. The basic element of the BACnet control system is a sub-station DESIGO PX PXC100-E.D. The DESIGO PX sub-stations are controlled by a user-friendly control panel PXM 20E [ 43 ]. The BACnet technology is used for HVAC (heating, ventilating, and air conditioning) control in SHC. The process sub-station BACnet/IP DESIGO PX-PXC100 E.D. forms a foundation of the control system for the measurement of nonelectrical quantities and control of operating and technical functions in SH [ 44 ]. The communication between the individual modules is executed via a standard BACnet protocol over Ethernet. Mutual communication between the sub-stations (peer-to-peer) is supported there [ 45 ]. KNX technology is used for lighting and blinds control and the switching on/off of socket circuits [ 46 ]. Furthermore, information about the security status of the building is transferred from the Electronic Security System (ESS) using KNX technology [ 47 ]. The KNX technology is suitably integrated into the BACnet technology using an interface providing interoperability between the communication protocols [ 48 ]. The visualization of control and regulation of the operating and technical functions in the building Sensors 2019,19, 1407 4 of 28 (HVAC, blinds, lighting, etc.) and storage of the measured data in the database was created using DESIGO Insight software tools [ 49 ]. To unify the monitoring and control of various automation and electronic systems into a single environment, it is necessary to use a software tool that allows you to consolidate several types of bus systems, standards and communication, and data protocols into o single monitoring application. In our case, we took advantage of the opportunities offered by the PI System software tool (hereinafter referred to as PI) produced by OSIsoft (Figure 1) [50]. Sensors 2019, 19, 1407 4 of 27 a software tool that allows you to consolidate several types of bus systems, standards and communication, and data protocols into o single monitoring application. In our case, we took advantage of the opportunities offered by the PI System software tool (hereinafter referred to as PI) produced by OSIsoft (Figure 1) [50]. Figure 1. A block diagram of data transfer from the Smart Home Care technology through the PI OPC (Ole for Process Control) interface to the PI Server and PI ProcessBook. The PI System includes SW tools such as PI ProcessBook for user-friendly data readout with the ability to create an application for the visualization and monitoring of SHC resident activities (Figures 2 and 3) [51–53]. Figure 2. The main visualization screen in SW tool PI ProcessBook, first floor. Figure 1. A block diagram of data transfer from the Smart Home Care technology through the PI OPC (Ole for Process Control) interface to the PI Server and PI ProcessBook. The PI System includes SW tools such as PI ProcessBook for user-friendly data readout with the ability to create an application for the visualization and monitoring of SHC resident activities (Figures 2and 3) [51–53]. Sensors 2019, 19, 1407 4 of 27 a software tool that allows you to consolidate several types of bus systems, standards and communication, and data protocols into o single monitoring application. In our case, we took advantage of the opportunities offered by the PI System software tool (hereinafter referred to as PI) produced by OSIsoft (Figure 1) [50]. Figure 1. A block diagram of data transfer from the Smart Home Care technology through the PI OPC (Ole for Process Control) interface to the PI Server and PI ProcessBook. The PI System includes SW tools such as PI ProcessBook for user-friendly data readout with the ability to create an application for the visualization and monitoring of SHC resident activities (Figures 2 and 3) [51–53]. Figure 2. The main visualization screen in SW tool PI ProcessBook, first floor. Figure 2. The main visualization screen in SW tool PI ProcessBook, first floor. Sensors 2019,19, 1407 5 of 28 Sensors 2019, 19, 1407 5 of 27 Figure 3. A screen for room 204 with a detailed description of individual technologies and individual charts in SW tool PI ProcessBook. Visualization for the SHC Created in the PI ProcessBook Tool Visualization of the wooden house in the PI Process Book tool is divided into several screens, which can be continuously accessed from the main screen (Figure 2). The SHC control technology is integrated on each visualization screen in accordance with the Building Management System (BMS). These screens further comprise of buttons for entry into individual rooms. After clicking on the relevant room button, a new window, containing a detailed description of the technology used in the specific room shown in the individual charts, will appear. Each technological element is illustrated in the chart, wherein the individual charts are sorted in accordance with the groups of elements used (Figure 3). The individual technology units are then displayed on separate screens and can also be viewed from the main screen. This solution was chosen because it is not possible to place all the information about the technologies implemented on one screen so that the screen remained wellarranged. The additional distribution of the technologies into individual screens will allow the user to get a better insight into what elements belong to the individual technologies and which do not anymore. Thanks to this solution selected, orientation in the enclosed charts are easier as well as the analysis of the individual quantities and actions in the building. The visualization, monitoring, and processing of the measured values of non-electric variables, such as the measurement of temperature, humidity, and CO 2 for monitoring the quality of the indoor environment of the selected room in the building described, are implemented using the PI System software application and SPSS IBM SW Tool (Figure 1). 3. Proposed Method for Creating an Optimized Model of CO 2 Concentration Prediction 3.1. Implementation of Predictive Analysis Using the IBM SPSS Modeler The IBM SPSS Modeler allows users to build models using a simplified, easy-to-use, and objectoriented user interface. The user is provided with Modelling algorithms, such as prediction, classification, segmentation, and association detection. The model results can be easily deployed and read into databases, IBM SPSS Statistics and a wide variety of other applications. Working with IBM SPSS Modeler can be divided into three basic steps: Figure 3. A screen for room 204 with a detailed description of individual technologies and individual charts in SW tool PI ProcessBook. Visualization for the SHC Created in the PI ProcessBook Tool Visualization of the wooden house in the PI Process Book tool is divided into several screens, which can be continuously accessed from the main screen (Figure 2). The SHC control technology is integrated on each visualization screen in accordance with the Building Management System (BMS). These screens further comprise of buttons for entry into individual rooms. After clicking on the relevant room button, a new window, containing a detailed description of the technology used in the specific room shown in the individual charts, will appear. Each technological element is illustrated in the chart, wherein the individual charts are sorted in accordance with the groups of elements used (Figure 3). The individual technology units are then displayed on separate screens and can also be viewed from the main screen. This solution was chosen because it is not possible to place all the information about the technologies implemented on one screen so that the screen remained well-arranged. The additional distribution of the technologies into individual screens will allow the user to get a better insight into what elements belong to the individual technologies and which do not anymore. Thanks to this solution selected, orientation in the enclosed charts are easier as well as the analysis of the individual quantities and actions in the building. The visualization, monitoring, and processing of the measured values of non-electric variables, such as the measurement of temperature, humidity, and CO 2 for monitoring the quality of the indoor environment of the selected room in the building described, are implemented using the PI System software application and SPSS IBM SW Tool (Figure 1). 3. Proposed Method for Creating an Optimized Model of CO2Concentration Prediction 3.1. Implementation of Predictive Analysis Using the IBM SPSS Modeler The IBM SPSS Modeler allows users to build models using a simplified, easy-to-use, and object-oriented user interface. The user is provided with Modelling algorithms, such as prediction, classification, segmentation, and association detection. The model results can be easily deployed and read into databases, IBM SPSS Statistics and a wide variety of other applications. Working with IBM SPSS Modeler can be divided into three basic steps: Sensors 2019,19, 1407 6 of 28 1. Importing the data into IBM SPSS Modeler 2. Performing a series of analyses on the imported data 3. Evaluation and exporting the data This sequence is also known as a Datastream because the data is flowing from the source to each analysis node and then to the output. The IBM SPSS Modeler allows the users to work with multiple data streams at once. These data streams can be build and modified using the stream canvas area of the application. These streams are created by drawing diagrams of relevant data operations. IBM SPSS Modeler’s Node Palette area of the displays shows most of the available data and modeling tools. The user may perform a simple drag and drop on each item in the nodes palette to add them to the current stream. The node palette items are divided into a few main categories as follows [54]: • Source: contains nodes that allow importing data into IBM SPSS Modeler from external sources such as analytic servers, databases, XML files, Microsoft Excel etc. • Record Ops: includes Nodes performing operations on data records, such as selecting, merging, and appending. • Field Ops: these nodes can perform operations on data fields such as filtering, deriving new fields, and determining the measurement level for given fields. •Graphs: Provides nodes that can graphically represent the data from before and after modeling. • Modeling: contains available modeling algorithms such as neural networks, decision trees, clustering, and data sequencing. •Output: consists of nodes that can provide outputs, such as plots, charts, evaluation, etc. • Export: composed of the nodes that can export the output to other applications, such as Microsoft Excel. •IBM SPSS Statistics: dedicated to nodes for importing or exporting data to IBM SPSS Statistics. Neural networks are one of the many ways to achieve predictive analysis. IBM SPSS Modeler offers multiple types of neural networks for predictive analysis. The text further describes the procedure for determining the appropriate method of predicting the course of CO 2 concentration from the measured values taken by the indoor temperature sensor T i ( ◦ C), (QPA 2062) in an SHC room (range 0 to 50 ◦ C/ − 35 to 35 ◦ C, accuracy ± 1K) and relative humidity rH (%), (QPA 2062) (range 0 and 100%, accuracy ± 5%), outdoor temperature T o ( ◦ C), (AP 257/22), (range: − 30 . . . + 80 ◦ C, resolution: 0.1 ◦ C) using the RBF. The RBF was selected due to its higher speed of training [ 55 ]. The RBF network is a feed-forward network that requires supervised learning. Unlike multilayer perceptron’s (MLP), this network consists of only one hidden layer. Overall, there are three layers in the RBF network: the input layer, RBF layer, and the output layer. The IBM SPSS algorithm guide describes mathematical models of these layers [56] (Figure 4) as following: Input layer:J0=P units,a0:1, . . . , a0:J0with a0:j=Xj RBF layer:j1units,units,a1:1, . . . a1:j1; with a1:j=∅j(X) ∅j(X) = e (−∑P p=11 2σ2 jp (xp−µjp)2) ∑J1 j=1e (−∑P p=11 2σ2 jp (xp−µjp)2)(1) Output layer:j2=R units,aI:1, . . . aI:j2with aI:r=ωI:r+∑j1 j=1ωrj∅j(X) Where: X(m): Input vector I: Number of layers (for RBM = 2) Ji: Number of units in layer i ∅j(X(m)): jth unit for input X(m),j=1, . . . , j1. Sensors 2019,19, 1407 7 of 28 µj: Center of ∅j σj: Width of ∅j am i:j: Unit j of layer i ωrj: weight connecting rth output unit and jth hidden unit of RBF layer Sensors 2019, 19, 1407 7 of 27 𝑎: : Unit j of layer i 𝜔: weight connecting rth output unit and jth hidden unit of RBF layer Figure 4. The Radial Basis Function (RBF) neural network diagram. The training of RBF can be divided into two stages. The first stage determines the basis function by clustering methods and the second stage determines the weights given to the basis function. SPSS measures the accuracy of neural networks by calculating the percentage of the records for which the predicted value matches the observed value. For the continues values, the accuracy is calculated by 1 minus the average of the absolute values of the predicted values minus the observed values over the maximum predicted value minus the minimum predicted value (the following formula) [56]. Accuracy =  ∑(1− ()  ()󰇻 ()(())   ), (2) 3.2. Signal Trend Detection Based on Wavelet Transformation Materials and Methods In this section, we introduce a method for the CO2 concentration prediction optimization based on the Wavelet transformation additive noise canceling. Based on the experimental results, the predicted CO2 trend contains glitches representing the fast change part of the signal. Such signal segments may significantly deteriorate the quality of the prediction. We propose an optimized scheme of the neural network prediction based on the Wavelet filtration appearing as a robust method due to a wide variability of the filtration settings. Such a system significantly improves the prediction system based on the neural network. In the signal processing, we assume that each signal y(t) is composed of two essential parts, namely, they are the signal trend T(t) and a component having a stochastic character X(t) which is perceived as the signal noise and details. Based on this definition, we can use the following signal formulation (3): 𝑦(𝑡)=𝑇(𝑡)+𝑋(𝑡) (3) The major problem when the signal trend is being extracted is noise detection. There are many applications of the trend detection including the CO2 measurement. Such a signal may be influenced by the glitches which should be removed to obtain a smooth signal for further processing. The wavelet analysis represents a transformation of the signal y(t) to obtain two types of coefficients, particularly they are the wavelet and scaling coefficients. These coefficients are completely equivalent with the original CO2 signal. It is supposed that wavelet coefficients are related to changes along a specifically defined scale. The main idea of the signal trend detection is to perform an association of the scaling coefficients with the signal trend T(x). On the other hand, the wavelet coefficients are supposed to be associated with the signal noise, which is mainly represented by the glitches when processing the CO2 signal. In our analysis, we considered an uncorrelated noise, adapting the wavelet estimator to work as a kernel estimator. The advantage of such an approach is formulating an Figure 4. The Radial Basis Function (RBF) neural network diagram. The training of RBF can be divided into two stages. The first stage determines the basis function by clustering methods and the second stage determines the weights given to the basis function. SPSS measures the accuracy of neural networks by calculating the percentage of the records for which the predicted value matches the observed value. For the continues values, the accuracy is calculated by 1 minus the average of the absolute values of the predicted values minus the observed values over the maximum predicted value minus the minimum predicted value (the following formula) [56]. Accuracy =1 n M ∑ m=1 (1−|yr(m)−ˆ yr(m)| maxm(yr(m))−minm(yr(m))), (2) 3.2. Signal Trend Detection Based on Wavelet Transformation Materials and Methods In this section, we introduce a method for the CO 2 concentration prediction optimization based on the Wavelet transformation additive noise canceling. Based on the experimental results, the predicted CO 2 trend contains glitches representing the fast change part of the signal. Such signal segments may significantly deteriorate the quality of the prediction. We propose an optimized scheme of the neural network prediction based on the Wavelet filtration appearing as a robust method due to a wide variability of the filtration settings. Such a system significantly improves the prediction system based on the neural network. In the signal processing, we assume that each signal y(t) is composed of two essential parts, namely, they are the signal trend T(t) and a component having a stochastic character X(t) which is perceived as the signal noise and details. Based on this definition, we can use the following signal formulation (3): y(t)=T(t)+X(t)(3) The major problem when the signal trend is being extracted is noise detection. There are many applications of the trend detection including the CO 2 measurement. Such a signal may be influenced by the glitches which should be removed to obtain a smooth signal for further processing. The wavelet analysis represents a transformation of the signal y(t) to obtain two types of coefficients, particularly they are the wavelet and scaling coefficients. These coefficients are completely equivalent with the Sensors 2019,19, 1407 8 of 28 original CO 2 signal. It is supposed that wavelet coefficients are related to changes along a specifically defined scale. The main idea of the signal trend detection is to perform an association of the scaling coefficients with the signal trend T(x). On the other hand, the wavelet coefficients are supposed to be associated with the signal noise, which is mainly represented by the glitches when processing the CO 2 signal. In our analysis, we considered an uncorrelated noise, adapting the wavelet estimator to work as a kernel estimator. The advantage of such an approach is formulating an estimator based on the sampled data irregularity. In this method, we used the scaling coefficients as estimators of the signal trend. We supposed that the sampled CO 2 observations are represented by Y(tn) , thus, the CO 2 estimator is given by Equation (4): ˆ T(t) = N−1 ∑ n=0 Y(tn)ZEJ(t,s)ds (4) Integration is done over a set of the intervals ( An(s) ), their union forms perform partitioning interval covering all the observations tn, where tn∈An. Consequently, EJis defined as Equation (5): EJ(t,s) = 2−J∑ k∈Z θ(2−Jt−k)θ(2−Js−k)(5) In this expression, θ(t) represents the scaling function. This function is defined as follows (Equation (6)): θ(t) = ∑ k∈Z ckθ(2t−k)(6) The wavelet function is defined by Equation (7): ψ(t) = (−1)kc1−kθ(2t−k)(7) The first crucial task is an appropriate selection of the mother’s wavelet for the predicted CO 2 signal filtration. Supposing the Daubechies wavelets can well reflect morphological structure therefore, this family was used for our model. Particularly, in our approach, we used the Daubechies wavelet (Db6), with the D6 scaling function utilizing the orthogonal Daubechies coefficients. 3.3. Validation Ratings Used In order to carry out the objective comparison, the following parameters were considered: Mean Absolute Error (MAE) represents the estimator measuring of the difference between two continuous variables. The MAE is given by the following expression: MAE =1 n n ∑ i=1 |yi−ˆ yı|(8) Mean squared error (MSE) represents the estimator measuring the average of the error squares between two signals. The MSE represents a risk function which corresponds with the expected value of the squared or quadratic error loss. The MSE is given by the following expression: MSE(x1,ˆ x2) = 1 n n ∑ i=1 (x(i)−ˆ x(i))2(9) Euclidean distance (ED) represents an ordinary straight-line distance between two points lying in the Euclidean space. Based on this distance, the Euclidean space becomes a metric space. The Sensors 2019,19, 1407 9 of 28 lower the Euclidean distance we achieve, the more similar are two signal samples. In our analysis, we considered a mean of the ED. The Euclidean distance is given by the following expression: d(x1,ˆ x2) = q(x1−ˆ x2)2+ (y1−ˆ y2)2(10) City Block distance (CB) represents a distance between two signals x1 , ˆ x2 in the space with the Cartesian coordinate system. This parameter can be interpreted as a sum of the lengths of the projections of the line segments between the points onto the coordinate axes. CB distance is defined as follows: dcb(x1,ˆ x2) =kx1−ˆ x2k= n ∑ i=1 |x1i−ˆ x2i|(11) The Correlation coefficient (R) measures a level of the linear dependency between two signals. The more the signals are considered linearly dependable, the higher the correlation coefficient is. In comparison with the previous parameters, the correlation coefficient represents a normative parameter. Zero correlation stands for the total dissimilarity between two signals, measured in a sense of their linear dependency. Contrarily, 1 and −1 stand for full positive and full negative correlation. As we have already stated above, in our work, we analyze two-month CO 2 predictions. In each measurement, we have a prediction from the neural network with 10, 50, 100, 150, 200, 250, 300, 350 and 400 neurons. Thus, we completely analyzed 9 predicted signals for each measurement. These signals are compared against the reference based on the evaluation parameters stated above. In terms of the Euclidean distance and MSE, lower values indicate a higher agreement between the signal and reference and thus, a better result. Contrarily, a higher correlation coefficient shows better results. In the following part of the analysis, we report the results of the quantification comparison. All the testing is done for the Wavelet Db6, with 6-level decomposition and the Wavelet settings as follows: threshold selection rule—Stein’s Unbiased Risk and soft thresholding for selection of the detailed coefficients. 4. The Practical Experimental Part 4.1. First Part of the ADL Information in SHC from the CO2Concentration Course, Blinds, Slats and On/Off Control of Lights The first experimental part in the study addressed the real needs of seniors who live in their own flats despite advanced age and mental and physical disabilities. These people strive to maintain maximum self-sufficiency and, thus, remove as much burden from their relatives, neighbors, friends or surroundings as possible. An example is a married couple, one of whom is mentally impaired, the other being the caregiver. They stay in touch with their family (their children) by SMS to keep them informed about how they are. In situations of acute need, the children are ready to come and help. In this example, the indirect ADL (Activities of Daily Living) in Room R203 (Figure 2) in SHC can be detected by monitoring operational and technical functions, such as •turning the lights on/off (Figure 5) •opening/closing windows (Figure 6) •raising/lowering blinds (Figure 7) •rotating blind slats (Figure 8) •increasing/decreasing the CO2concentration (Figures 5–8). Sensors 2019,19, 1407 16 of 28 Sensors 2019, 19, 1407 16 of 27 3 100 82.4 0.925 7.29 × 10−3 6.75 × 10−04 4 150 89.8 0.927 9.29 × 10−3 1.18 × 10−03 5 200 91.2 0.935 8.15 × 10−3 1.02 × 10−03 6 250 91.1 0.931 9.14 × 10−3 1.45 × 10−03 7 300 91.2 0.935 7.71 × 10−3 1.15 × 10−03 8 350 90.2 0.936 7.71 × 10−3 1.15 × 10−03 9 400 92.1 0.936 7.85 × 10−3 1.17 × 10−03 Figure 11 clearly demonstrates the relationship between accuracy, the experiments interval length and the number of neurons. By summing up all of the obtained results, it is clear that model number 3 (Table 6) trained with the data from 15th of May holds the highest accuracy (99.8%), linear correlation (0.996) and relevantly small error values. Model number 9 shows the best overall average accuracy (Table 9) and, in case of the experiments with the interval of 12th (see Figure 12) and 15th May, the accuracy difference between this model and the most accurate model (model number 3) is negligible. Therefore, model number 9 trained with the May 15th interval was selected as the most suitable model for the next stages (Figure 12). Figure 11. The accuracy for each interval of experiments. Figure 12. Model number 9 trained and validated using data from yjr 15th of May 2018. Figure 11. The accuracy for each interval of experiments. Sensors 2019, 19, 1407 16 of 27 3 100 82.4 0.925 7.29 × 10−3 6.75 × 10−04 4 150 89.8 0.927 9.29 × 10−3 1.18 × 10−03 5 200 91.2 0.935 8.15 × 10−3 1.02 × 10−03 6 250 91.1 0.931 9.14 × 10−3 1.45 × 10−03 7 300 91.2 0.935 7.71 × 10−3 1.15 × 10−03 8 350 90.2 0.936 7.71 × 10−3 1.15 × 10−03 9 400 92.1 0.936 7.85 × 10−3 1.17 × 10−03 Figure 11 clearly demonstrates the relationship between accuracy, the experiments interval length and the number of neurons. By summing up all of the obtained results, it is clear that model number 3 (Table 6) trained with the data from 15th of May holds the highest accuracy (99.8%), linear correlation (0.996) and relevantly small error values. Model number 9 shows the best overall average accuracy (Table 9) and, in case of the experiments with the interval of 12th (see Figure 12) and 15th May, the accuracy difference between this model and the most accurate model (model number 3) is negligible. Therefore, model number 9 trained with the May 15th interval was selected as the most suitable model for the next stages (Figure 12). Figure 11. The accuracy for each interval of experiments. Figure 12. Model number 9 trained and validated using data from yjr 15th of May 2018. Figure 12. Model number 9 trained and validated using data from yjr 15th of May 2018. Sensors 2019, 19, 1407 17 of 27 Figure 13. Model number 9 trained and validated using data from the interval of the 6 th of May 2018 to 13 th of May 2018. Figure 14. Model number 3 trained and validated using data from the interval of 12 th of May 2018. 4.2.5. IoT Implementation with Watson Studio SW Tool The data streams created in the IBM SPSS Modelers can be stored as a file (“.str” format). The IBM Watson Studio allows the user to import the developed data stream simply by uploading the stored files. This allows the data streams that were originally developed in the IBM SPSS Modeler to take advantage of cloud computing, cloud storage and the possibility of near-real-time streaming. As it was explained earlier, model number 9 (with 400 neurons) trained with the data from 15th of May was selected as the best overall result of this experiment. Specifically, this model showed high accuracy, high linear correlation, low MAE and MSE errors. Therefore, it was uploaded to Watson studio. Figure 15 shows the streamed developed in SPSS in Watson studio for near-real-time training (Excel files were replaced with assets on the cloud). Figure 16 shows a data flow stream in Watson that includes model 9 trained with data from the 15 th of May for near real-time prediction. Figure 15. Watson Studio with the data stream developed in IBM SPSS Modeler. Figure 13. Model number 9 trained and validated using data from the interval of the 6 th of May 2018 to 13th of May 2018. 4.2.4. Analyzing the Results and Selecting the Best Model Table 8shows the average of the accuracy, linear correlation MAE and MSE in each experiment. By observing this table, it is apparent that the experiments with an interval length of one month, hold the lowest average accuracies (63.1% and 72%). It can also be observed that the experiment with the 15 th of May as period holds the highest average accuracy (99.5%), highest linear correlation (0.995) and relevantly low error values (MAE: 1.78 × 10 −3 , MSE: 2.44 × 10 −5 ). Additionally, the Sensors 2019,19, 1407 17 of 28 experiment with a interval length of a week in May shows a slightly smaller average accuracy (94.7%) and linear correlation (0.967) but it has overall lower error values (MAE: 1.78 × 10 −3 , MSE: 2.44 × 10 −5 ) (Figure 14). Table 8. The average of the obtained results in each experiment. Measurement Period Accuracy (%) Validation Linear Correlation MAE MSE May 63.1 0.78 7.44 ×10−31.51 ×10−4 November 72.2 0.848 1.30 ×10−21.45 ×10−3 May 6 to 13 94.7 0.967 1.33 ×10−37.52 ×10−6 November 6 to 13 89.9 0.939 5.00 ×10−31.17 ×10−4 May 12 95.7 0.940 3.12 ×10−26.14 ×10−3 May 15 99.5 0.995 1.78 ×10−32.44 ×10−5 November 15 92.0 0.912 3.22 ×10−39.67 ×10−5 Sensors 2019, 19, 1407 17 of 27 Figure 13. Model number 9 trained and validated using data from the interval of the 6 th of May 2018 to 13 th of May 2018. Figure 14. Model number 3 trained and validated using data from the interval of 12 th of May 2018. 4.2.5. IoT Implementation with Watson Studio SW Tool The data streams created in the IBM SPSS Modelers can be stored as a file (“.str” format). The IBM Watson Studio allows the user to import the developed data stream simply by uploading the stored files. This allows the data streams that were originally developed in the IBM SPSS Modeler to take advantage of cloud computing, cloud storage and the possibility of near-real-time streaming. As it was explained earlier, model number 9 (with 400 neurons) trained with the data from 15th of May was selected as the best overall result of this experiment. Specifically, this model showed high accuracy, high linear correlation, low MAE and MSE errors. Therefore, it was uploaded to Watson studio. Figure 15 shows the streamed developed in SPSS in Watson studio for near-real-time training (Excel files were replaced with assets on the cloud). Figure 16 shows a data flow stream in Watson that includes model 9 trained with data from the 15 th of May for near real-time prediction. Figure 15. Watson Studio with the data stream developed in IBM SPSS Modeler. Figure 14. Model number 3 trained and validated using data from the interval of 12th of May 2018. Table 9contains the average results from all experiments for each model. The overall trend of this Table points toward the fact that with the increasing number of neurons, the accuracy of the model’s increases. This increase may be insignificant after a certain point. For example, models number 5, 6, 7, 8 and 9 share a similar average accuracy (91.2%, 91.1%, 90.2%, 90.2% and 92.1%), linear correlation (0.935, 0.931, 0.935, 0.936 and 0.936), MAE (8.15 × 10 −3 , 9.14 × 10 −3 , 7.71 × 10 −3 , 7.71 × 10 −3 and 7.85 ×10−3) and MSE (1.02 ×10−3, 1.45 ×10−3, 1.15 ×10−3, 1.15 ×10−3and 1.17 ×10−3). Table 9. The average of the obtained results in the six performed experiments. Order of Measurement (Model Number) Number of Neurons (-) Accuracy (%) Validation Linear Correlation MAE MSE 1 10 64.0 0.774 1.94 ×10−21.83 ×10−3 2 50 87.8 0.899 8.88 ×10−36.38 ×10−4 3 100 82.4 0.925 7.29 ×10−36.75 ×10−4 4 150 89.8 0.927 9.29 ×10−31.18 ×10−3 5 200 91.2 0.935 8.15 ×10−31.02 ×10−3 6 250 91.1 0.931 9.14 ×10−31.45 ×10−3 7 300 91.2 0.935 7.71 ×10−31.15 ×10−3 8 350 90.2 0.936 7.71 ×10−31.15 ×10−3 9 400 92.1 0.936 7.85 ×10−31.17 ×10−3 Figure 11 clearly demonstrates the relationship between accuracy, the experiments interval length and the number of neurons. By summing up all of the obtained results, it is clear that model number 3 (Table 6) trained with the data from 15 th of May holds the highest accuracy (99.8%), linear correlation (0.996) and relevantly small error values. Model number 9 shows the best overall average accuracy Sensors 2019,19, 1407 18 of 28 (Table 9) and, in case of the experiments with the interval of 12 th (see Figure 12) and 15 th May, the accuracy difference between this model and the most accurate model (model number 3) is negligible. Therefore, model number 9 trained with the May 15 th interval was selected as the most suitable model for the next stages (Figure 12). 4.2.5. IoT Implementation with Watson Studio SW Tool The data streams created in the IBM SPSS Modelers can be stored as a file (“.str” format). The IBM Watson Studio allows the user to import the developed data stream simply by uploading the stored files. This allows the data streams that were originally developed in the IBM SPSS Modeler to take advantage of cloud computing, cloud storage and the possibility of near-real-time streaming. As it was explained earlier, model number 9 (with 400 neurons) trained with the data from 15th of May was selected as the best overall result of this experiment. Specifically, this model showed high accuracy, high linear correlation, low MAE and MSE errors. Therefore, it was uploaded to Watson studio. Figure 15 shows the streamed developed in SPSS in Watson studio for near-real-time training (Excel files were replaced with assets on the cloud). Figure 16 shows a data flow stream in Watson that includes model 9 trained with data from the 15th of May for near real-time prediction. Sensors 2019, 19, 1407 17 of 27 Figure 13. Model number 9 trained and validated using data from the interval of the 6 th of May 2018 to 13 th of May 2018. Figure 14. Model number 3 trained and validated using data from the interval of 12 th of May 2018. 4.2.5. IoT Implementation with Watson Studio SW Tool The data streams created in the IBM SPSS Modelers can be stored as a file (“.str” format). The IBM Watson Studio allows the user to import the developed data stream simply by uploading the stored files. This allows the data streams that were originally developed in the IBM SPSS Modeler to take advantage of cloud computing, cloud storage and the possibility of near-real-time streaming. As it was explained earlier, model number 9 (with 400 neurons) trained with the data from 15th of May was selected as the best overall result of this experiment. Specifically, this model showed high accuracy, high linear correlation, low MAE and MSE errors. Therefore, it was uploaded to Watson studio. Figure 15 shows the streamed developed in SPSS in Watson studio for near-real-time training (Excel files were replaced with assets on the cloud). Figure 16 shows a data flow stream in Watson that includes model 9 trained with data from the 15 th of May for near real-time prediction. Figure 15. Watson Studio with the data stream developed in IBM SPSS Modeler. Figure 15. Watson Studio with the data stream developed in IBM SPSS Modeler. Sensors 2019, 19, 1407 18 of 27 Figure 16. Watson Studio with the data flow stream that uses model 9. 4.2.6. Discussion of the Second Experimental Part By evaluating the obtained results from the implementation with IBM SPSS Modeler, it is apparent that as it was expected that the experiments with interval length of one day showed better overall accuracy (average value up to 99.5%) and the experiments with sample periods of one month showed the least overall accuracy (average value up to 72.2%). Additionally, in four out of six experiments, model number 9 held the highest accuracy, and in the other two cases, it had the only insignificant difference with the most accurate models. Therefore, it was selected as the overall most accurate model. As it was mentioned earlier, in terms of the training interval, 15 th of May showed the most accurate results. Therefore, model number 9 with the 15 th of May training interval (Table 9) was selected and exported to an IBM cloud data stream. Furthermore, the results demanded additional filtering in order to reduce the noise and provide smoother results. 4.3. Third Part—Testing and Quantitative Comparison (WT Additive Noise Canceling) 4.3.1. Testing and Quantitative Comparison In our research, we analyzed signals representing the CO 2 signals. We had a set of the estimated (predicted) signals being compared against the real measured signal CO 2 , which is perceived as a reference. In our analysis, we are comparing two-month CO 2 prediction. We compared one-day, oneweek and one-month predictions for May and November 2018. Based on the observations, it is apparent that the predicted CO 2 signals do not have a smooth process. They are frequently influenced by rapid oscillations, so-called glitches and signal fluctuations (Figures 17 and 18). Such signal variations represent the signal noise, impairing the real trend of the CO 2 prediction, which should be reduced. In our analysis, we used wavelet filtration to eliminate such signals to obtain the signal trend for further processing. As we have already stated above, we used the mother’s wavelet Db6 for the CO 2 signal trend detection. Firstly, we take advantage of the fact that different level of the decomposition allows perceiving more or less signal details represented by the detailed coefficients. Since we need to perceive the signal trend by eliminating the steep fluctuations, we need to consider an appropriate level of the decomposition. An experimental comparison of individual wavelet settings is reported in Figures 17 and 18. Based on the experimental results, we used the 6-level decomposition for the CO 2 signal trend detection. The filtration procedure further utilizes the following settings: threshold selection rule—Stein’s Unbiased Risk and soft thresholding for selection of the detailed coefficients. Figure 17. The comparison of the original predicted CO 2 signal and Wavelet filtration from 15 May 2018 for 2-level decomposition (left), 6-level decomposition (middle) and 7-level decomposition (right). Figure 16. Watson Studio with the data flow stream that uses model 9. 4.2.6. Discussion of the Second Experimental Part By evaluating the obtained results from the implementation with IBM SPSS Modeler, it is apparent that as it was expected that the experiments with interval length of one day showed better overall accuracy (average value up to 99.5%) and the experiments with sample periods of one month showed the least overall accuracy (average value up to 72.2%). Additionally, in four out of six experiments, model number 9 held the highest accuracy, and in the other two cases, it had the only insignificant difference with the most accurate models. Therefore, it was selected as the overall most accurate model. As it was mentioned earlier, in terms of the training interval, 15 th of May showed the most accurate results. Therefore, model number 9 with the 15 th of May training interval (Table 9) was selected and exported to an IBM cloud data stream. Furthermore, the results demanded additional filtering in order to reduce the noise and provide smoother results. 4.3. Third Part—Testing and Quantitative Comparison (WT Additive Noise Canceling) 4.3.1. Testing and Quantitative Comparison In our research, we analyzed signals representing the CO 2 signals. We had a set of the estimated (predicted) signals being compared against the real measured signal CO 2 , which is perceived as a Sensors 2019,19, 1407 19 of 28 reference. In our analysis, we are comparing two-month CO 2 prediction. We compared one-day, one-week and one-month predictions for May and November 2018. Based on the observations, it is apparent that the predicted CO 2 signals do not have a smooth process. They are frequently influenced by rapid oscillations, so-called glitches and signal fluctuations (Figures 17 and 18). Such signal variations represent the signal noise, impairing the real trend of the CO 2 prediction, which should be reduced. In our analysis, we used wavelet filtration to eliminate such signals to obtain the signal trend for further processing. As we have already stated above, we used the mother’s wavelet Db6 for the CO 2 signal trend detection. Firstly, we take advantage of the fact that different level of the decomposition allows perceiving more or less signal details represented by the detailed coefficients. Since we need to perceive the signal trend by eliminating the steep fluctuations, we need to consider an appropriate level of the decomposition. An experimental comparison of individual wavelet settings is reported in Figures 17 and 18. Based on the experimental results, we used the 6-level decomposition for the CO 2 signal trend detection. The filtration procedure further utilizes the following settings: threshold selection rule—Stein’s Unbiased Risk and soft thresholding for selection of the detailed coefficients. Sensors 2019, 19, 1407 18 of 27 Figure 16. Watson Studio with the data flow stream that uses model 9. 4.2.6. Discussion of the Second Experimental Part By evaluating the obtained results from the implementation with IBM SPSS Modeler, it is apparent that as it was expected that the experiments with interval length of one day showed better overall accuracy (average value up to 99.5%) and the experiments with sample periods of one month showed the least overall accuracy (average value up to 72.2%). Additionally, in four out of six experiments, model number 9 held the highest accuracy, and in the other two cases, it had the only insignificant difference with the most accurate models. Therefore, it was selected as the overall most accurate model. As it was mentioned earlier, in terms of the training interval, 15 th of May showed the most accurate results. Therefore, model number 9 with the 15 th of May training interval (Table 9) was selected and exported to an IBM cloud data stream. Furthermore, the results demanded additional filtering in order to reduce the noise and provide smoother results. 4.3. Third Part—Testing and Quantitative Comparison (WT Additive Noise Canceling) 4.3.1. Testing and Quantitative Comparison In our research, we analyzed signals representing the CO 2 signals. We had a set of the estimated (predicted) signals being compared against the real measured signal CO 2 , which is perceived as a reference. In our analysis, we are comparing two-month CO 2 prediction. We compared one-day, oneweek and one-month predictions for May and November 2018. Based on the observations, it is apparent that the predicted CO 2 signals do not have a smooth process. They are frequently influenced by rapid oscillations, so-called glitches and signal fluctuations (Figures 17 and 18). Such signal variations represent the signal noise, impairing the real trend of the CO 2 prediction, which should be reduced. In our analysis, we used wavelet filtration to eliminate such signals to obtain the signal trend for further processing. As we have already stated above, we used the mother’s wavelet Db6 for the CO 2 signal trend detection. Firstly, we take advantage of the fact that different level of the decomposition allows perceiving more or less signal details represented by the detailed coefficients. Since we need to perceive the signal trend by eliminating the steep fluctuations, we need to consider an appropriate level of the decomposition. An experimental comparison of individual wavelet settings is reported in Figures 17 and 18. Based on the experimental results, we used the 6-level decomposition for the CO 2 signal trend detection. The filtration procedure further utilizes the following settings: threshold selection rule—Stein’s Unbiased Risk and soft thresholding for selection of the detailed coefficients. Figure 17. The comparison of the original predicted CO 2 signal and Wavelet filtration from 15 May 2018 for 2-level decomposition (left), 6-level decomposition (middle) and 7-level decomposition (right). Figure 17. The comparison of the original predicted CO 2 signal and Wavelet filtration from 15 May 2018 for 2-level decomposition (left), 6-level decomposition (middle) and 7-level decomposition (right). Sensors 2019, 19, 1407 19 of 27 Figure 18. The comparison of the original predicted CO 2 signal and Wavelet filtration from 1–29 May 2018 for 2-level decomposition (left), 6-level decomposition (middle) and 7-level decomposition (right). Wavelet filtration was used for the extraction of the CO 2 signal trend, simultaneously rapid changes of the signal were removed. On Figures 19 and 20, there is a comparison among the reference signal and predicted signals by wavelet transformation for day and month predictions from May and November 2018. Figure 19. An example of one-day prediction from 15 th May 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Figure 20. An example of whole-month prediction from November 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Figure 18. The comparison of the original predicted CO 2 signal and Wavelet filtration from 1–29 May 2018 for 2-level decomposition ( left ), 6-level decomposition ( middle ) and 7-level decomposition (right). Wavelet filtration was used for the extraction of the CO 2 signal trend, simultaneously rapid changes of the signal were removed. On Figures 19 and 20, there is a comparison among the reference signal and predicted signals by wavelet transformation for day and month predictions from May and November 2018. Sensors 2019,19, 1407 20 of 28 Sensors 2019, 19, 1407 19 of 27 Figure 18. The comparison of the original predicted CO 2 signal and Wavelet filtration from 1–29 May 2018 for 2-level decomposition (left), 6-level decomposition (middle) and 7-level decomposition (right). Wavelet filtration was used for the extraction of the CO 2 signal trend, simultaneously rapid changes of the signal were removed. On Figures 19 and 20, there is a comparison among the reference signal and predicted signals by wavelet transformation for day and month predictions from May and November 2018. Figure 19. An example of one-day prediction from 15 th May 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Figure 20. An example of whole-month prediction from November 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Figure 19. An example of one-day prediction from 15 th May 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Sensors 2019, 19, 1407 19 of 27 Figure 18. The comparison of the original predicted CO 2 signal and Wavelet filtration from 1–29 May 2018 for 2-level decomposition (left), 6-level decomposition (middle) and 7-level decomposition (right). Wavelet filtration was used for the extraction of the CO 2 signal trend, simultaneously rapid changes of the signal were removed. On Figures 19 and 20, there is a comparison among the reference signal and predicted signals by wavelet transformation for day and month predictions from May and November 2018. Figure 19. An example of one-day prediction from 15 th May 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Figure 20. An example of whole-month prediction from November 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Figure 20. An example of whole-month prediction from November 2018 of the CO 2 signal. Filtration is done by using the Db6 wavelet with 6-level decomposition. Based on the results, wavelet filtration is capable of filtering rapid signal changes whilst preserving the signal trend. To justify this situation, we report the selected situations showing the glitches deteriorating a smooth signal trend, and a respective wavelet approximation largely reducing such signal parts (Figures 21–23). Among these cases, we mark the most significant glitches as green in the originally predicted signals to highlight the Wavelet smoothing effectivity. As it is obvious, the CO 2 prediction contains lots of significant occurrences represented by the glitches and spikes, significantly deteriorating the smoothness of the analyzed signal. Wavelet appears to be a reliable alternative for reduction of those parts of the signal. On the other hand, we are aware that trend detection, in some cases, reduces the peaks and thus, the original signal’s amplitude is reduced. Such situations are reported in Figures 21–23. Sensors 2019,19, 1407 21 of 28 Sensors 2019, 19, 1407 20 of 27 Based on the results, wavelet filtration is capable of filtering rapid signal changes whilst preserving the signal trend. To justify this situation, we report the selected situations showing the glitches deteriorating a smooth signal trend, and a respective wavelet approximation largely reducing such signal parts (Figures 21, 22 and 23). Among these cases, we mark the most significant glitches as green in the originally predicted signals to highlight the Wavelet smoothing effectivity. Figure 21. The 6–13 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Figure 22. The 15 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Figure 21. The 6–13 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Sensors 2019, 19, 1407 20 of 27 Based on the results, wavelet filtration is capable of filtering rapid signal changes whilst preserving the signal trend. To justify this situation, we report the selected situations showing the glitches deteriorating a smooth signal trend, and a respective wavelet approximation largely reducing such signal parts (Figures 21, 22 and 23). Among these cases, we mark the most significant glitches as green in the originally predicted signals to highlight the Wavelet smoothing effectivity. Figure 21. The 6–13 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Figure 22. The 15 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Figure 22. The 15 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Sensors 2019, 19, 1407 20 of 27 Based on the results, wavelet filtration is capable of filtering rapid signal changes whilst preserving the signal trend. To justify this situation, we report the selected situations showing the glitches deteriorating a smooth signal trend, and a respective wavelet approximation largely reducing such signal parts (Figures 21, 22 and 23). Among these cases, we mark the most significant glitches as green in the originally predicted signals to highlight the Wavelet smoothing effectivity. Figure 21. The 6–13 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Figure 22. The 15 November 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Figure 23. The 4 November–3 December 2018 prediction of the CO 2 signal containing several glitches and signal spikes, marked as green rectangle. Sensors 2019,19, 1407 22 of 28 In the last part of our analysis, the objective comparison is carried out. As we have already stated, we are comparing predicted CO 2 signals with signals being filtered out by the wavelet transformation. All the signals are compared against the reference CO2signals for day and week predictions. As we have already stated above, in our work we analyzed the two-month CO2 prediction. In each measurement, we have a prediction from a neural network with 10, 50, 100, 150, 200, 250, 300, 350 and 400 neurons. Thus, we completely analyzed the predicted signals of 9 models for each measurement. These signals are compared against the reference based on the evaluation parameters stated above. In terms of the Euclidean distance and MSE, lower values indicate a higher agreement between the signal and reference and thus, better result. Contrarily, a higher correlation coefficient indicates better results. In the following part of the analysis, we report the results of the quantification comparison. All the testing is done for the Wavelet Db6, with 6-level decomposition and the following Wavelet settings: threshold selection rule—Stein’s Unbiased Risk and soft thresholding for selection of the detailed coefficients. Figure 24 shows the MSE evaluation for CO 2 prediction in the following time intervals: (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Figure 25 shows the correlation coefficient evaluation for the CO 2 prediction in the following time intervals: (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Figure 26 shows the Euclidean distance evaluation for the CO 2 prediction in the following time intervals: (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Sensors 2019, 19, 1407 21 of 27 Figure 23. The 4 November–3 December 2018 prediction of the CO2 signal containing several glitches and signal spikes, marked as green rectangle. As it is obvious, the CO 2 prediction contains lots of significant occurrences represented by the glitches and spikes, significantly deteriorating the smoothness of the analyzed signal. Wavelet appears to be a reliable alternative for reduction of those parts of the signal. On the other hand, we are aware that trend detection, in some cases, reduces the peaks and thus, the original signal’s amplitude is reduced. Such situations are reported in Figures 21, 22 and 23. In the last part of our analysis, the objective comparison is carried out. As we have already stated, we are comparing predicted CO 2 signals with signals being filtered out by the wavelet transformation. All the signals are compared against the reference CO 2 signals for day and week predictions. As we have already stated above, in our work we analyzed the two-month CO2 prediction. In each measurement, we have a prediction from a neural network with 10, 50, 100, 150, 200, 250, 300, 350 and 400 neurons. Thus, we completely analyzed the predicted signals of 9 models for each measurement. These signals are compared against the reference based on the evaluation parameters stated above. In terms of the Euclidean distance and MSE, lower values indicate a higher agreement between the signal and reference and thus, better result. Contrarily, a higher correlation coefficient indicates better results. In the following part of the analysis, we report the results of the quantification comparison. All the testing is done for the Wavelet Db6, with 6-level decomposition and the following Wavelet settings: threshold selection rule—Stein’s Unbiased Risk and soft thresholding for selection of the detailed coefficients. Figure 24 shows the MSE evaluation for CO 2 prediction in the following time intervals: (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Figure 25 shows the correlation coefficient evaluation for the CO 2 prediction in the following time intervals: (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Figure 26 shows the Euclidean distance evaluation for the CO 2 prediction in the following time intervals: (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. (a) (b) (c) (d) (e) (f) Figure 24. The Mean Squared Error (MSE) evaluation for the CO2 prediction (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Figure 24. The Mean Squared Error (MSE) evaluation for the CO 2 prediction ( a ) 6–13 May 2018, ( b ) 15 May 2018, ( c ) 1–29 May 2018, ( d ) 6–13 November 2018, ( e ) 15 November 2018, ( f ) 4 November–3 December 2018. Sensors 2019,19, 1407 23 of 28 Sensors 2019, 19, 1407 22 of 27 (a) (b) (c) (d) (e) (f) Figure 25. The correlation coefficient evaluation for the CO 2 prediction (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Lastly, we summarize the achieved results of all the predicted signals. In Tables 10 and 11, we compare the individual parameters—Mean Square Error (MSE), Correlation index (Corr) and Euclidean distance (ED)—for individual CO 2 predictions from May and November 2018. Each parameter is averaged for all the predictions and the difference (diff) between the original prediction and Wavelet smoothing is evaluated (Table 12). (a) (b) (c) (d) (e) (f) Figure 26. The Euclidean distance evaluation for CO2 prediction (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Table 10. The summarization of the evaluation parameters of the original CO2 prediction. Date of Prediction MSE [-] Corr [%] ED [-] Figure 25. The correlation coefficient evaluation for the CO 2 prediction ( a ) 6–13 May 2018, ( b ) 15 May 2018, ( c ) 1–29 May 2018, ( d ) 6–13 November 2018, ( e ) 15 November 2018, ( f ) 4 November–3 December 2018. Sensors 2019, 19, 1407 22 of 27 (a) (b) (c) (d) (e) (f) Figure 25. The correlation coefficient evaluation for the CO 2 prediction (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Lastly, we summarize the achieved results of all the predicted signals. In Tables 10 and 11, we compare the individual parameters—Mean Square Error (MSE), Correlation index (Corr) and Euclidean distance (ED)—for individual CO 2 predictions from May and November 2018. Each parameter is averaged for all the predictions and the difference (diff) between the original prediction and Wavelet smoothing is evaluated (Table 12). (a) (b) (c) (d) (e) (f) Figure 26. The Euclidean distance evaluation for CO2 prediction (a) 6–13 May 2018, (b) 15 May 2018, (c) 1–29 May 2018, (d) 6–13 November 2018, (e) 15 November 2018, (f) 4 November–3 December 2018. Table 10. The summarization of the evaluation parameters of the original CO2 prediction. Date of Prediction MSE [-] Corr [%] ED [-] Figure 26. The Euclidean distance evaluation for CO 2 prediction ( a ) 6–13 May 2018, ( b ) 15 May 2018, ( c ) 1–29 May 2018, ( d ) 6–13 November 2018, ( e ) 15 November 2018, ( f ) 4 November–3 December 2018. Lastly, we summarize the achieved results of all the predicted signals. In Tables 10 and 11, we compare the individual parameters—Mean Square Error (MSE), Correlation index (Corr) and Euclidean distance (ED)—for individual CO 2 predictions from May and November 2018. Each parameter is averaged for all the predictions and the difference (diff) between the original prediction and Wavelet smoothing is evaluated (Table 12). Sensors 2019,19, 1407 24 of 28 Table 10. The summarization of the evaluation parameters of the original CO2prediction. Date of Prediction MSE [-] Corr [%] ED [-] 6–13 May 2018 7.577 ×10−697.0 0.266 15 May 2018 2.436 ×10−599.5 0.177 1–29 May 2018 1.505 ×10−477.9 2.439 6–13 November 2018 1.172 ×10−494.2 1.017 15 November 2018 9.761 ×10−593.3 0.366 4 November–3 December 2018 4.065 ×10−479.8 3.429 Table 11. The summarization of the evaluation parameters of the wavelet smoothing CO2prediction. Date of Prediction MSE [-] Corr [%] ED [-] 6–13 May 2018 6.442 ×10−698.1 0.238 15 May 2018 1.473 ×10−699.6 0.157 1–29 May 2018 1.360 ×10−480.6 2.292 6–13 November 2018 1.047 ×10−495.6 0.925 15 November 2018 7.061 ×10−595.7 0.315 4 November–3 December 2018 3.781 ×10−481.3 3.281 Table 12. The difference parameters of CO2prediction. Date of Prediction Diff MSE [-] Diff Corr [%] Diff ED [-] 6–13 May 2018 1.134 ×10−60.0049 0.0277 15 May 2018 4.886 ×10−68.524 ×10−40.0196 1–29 May 2018 1.443 ×10−50.027 0.147 6–13 November 2018 1.258 ×10−50.0079 0.092 15 November 2018 2.699 ×10−50.022 0.051 4 November–3 December 2018 2.842 ×10−50.015 0.147 4.3.2. Discussion of the Third Experimental Part As it is obvious, the predicted CO2signals contain lots of significant occurrences represented by the glitches and spikes, significantly deteriorating the smoothness of the analyzed signals. Such steep fluctuations may have a significant impact on CO 2 accuracy. Wavelet appears to be a reliable alternative for reduction of those parts of the signal. On the other hand, we are aware that trend detection, in some cases, reduces the peaks and thus, the original signal’s amplitude is reduced. In our work, we have studied the Daubechies wavelet family. These wavelets, as it is known, can well reflect the morphological structure of the signals. We are particularly using the Db6 wavelet for trend detection. Alternatively, we mention the comparison in Reference [ 50 ] of the CO 2 filtration based on the LMS algorithm. In this study, the authors employed adaptive filtration. The main limitation of this method is a necessity of the reference signal and a slow adaptation of the filtration procedure, as well as depending on the accuracy of the step size parameter µ calculation and inaccurate determination of the arrival and departure time of the person from the monitored area. Furthermore, the Wavelet filtration presented in this study achieves better results in a context of the objective comparison against the LMS filtration. Wavelet filtration has a much stronger potential for the CO 2 filtration due to a possibility of the application of a variety wavelets allowing for the extraction of specific morphological signal features in various decomposition levels and, thus, better optimize the CO 2 prediction. These facts predetermine wavelets to be a robust system for the CO2prediction enhancement. In the last part of our analysis, the objective comparison is carried out. As we have already stated, we compared originally measured CO 2 signals with predicted signals being filtered out by the wavelets. To carry out the objective comparison, the following parameters are considered: when considering the MSE, we get better results for wavelet trend detection. This means that we have minimized the difference between the gold standard and filtered signals. The correlation coefficient gives higher Sensors 2019,19, 1407 25 of 28 values for the predicted CO 2 signals. Regardless, we have achieved just slight differences. The reason might be caused by the fact that the trend detection largely omits higher peaks, therefore, the linear dependence for wavelet filtration is smaller when compared with the predicted signals. Using the Wavelet filtration leads to more accurate results against the predicted signals and signals are much more smoothed, not containing steep fluctuations. On the other hand, we are aware of a certain loss of the amplitude. Therefore, in the future, it would be worth investigating the frequency features of the CO2signals to objectively determine frequency modifications while filtering by the wavelets. 5. Conclusions The authors of the paper focused on designing a methodology that specifies a procedure for processing data measured by sensors in an SHC environment for the purpose of indirectly monitoring the presence of people in an SHC area through KNX and BACnet technologies commonly applied in building automation. This paper explores the possibilities of improving accuracy in CO2 predictions in SHC using IBM SPSS software tools in the IoT to determine the occupancy times of a monitored SHC room. The RBF method was applied to predict CO 2 levels from the measured indoor and outdoor temperatures and relative humidity. The accuracy of CO 2 predictions from the processed data was compared and evaluated at daily, weekly and monthly intervals for the spring and autumn periods. As it was expected the most accurate results were provided by experiments with the daily intervals (accuracy was about 99%) while the monthly intervals resulted in the least accurate results (accuracy was about 80%). Overall, the developed stream in IBM SPSS Modeler is capable of predicting the CO 2 concentration values using the values of humidity and indoor and outdoor temperature. By providing a live data asset to the IBM Cloud, the uploaded model can achieve near real-time prediction of CO 2 concentration values. Using a wavelet transform mathematical method to cancel additive noise led to more accurate results in predicted signals. The signals were also much smoother and did not contain sharp fluctuations, although there was a certain loss in amplitude, resulting in inaccuracies when the maximum achieved CO 2 value was determined. Future work should, therefore, focus on finding an optimal method for canceling additive noise in real time, which would help increase the overall accuracy of CO 2 predictions [ 57 – 60 ]. Additionally, the real-life and live performance of this implementation should be examined. Author Contributions: J.V. methodology; J.V., J.K. (Jan Kubicek), O.M.G. software; J.V., J.K. (Jan Kubicek) validation; J.V., J.K. (Jan Kubicek), O.M.G. formal analysis; J.V. investigation; J.V., O.M.G. resources; J.V. data curation; J.V., J.K. (Jan Kubicek), O.M.G. writing—original draft preparation; J.V., O.M.G. writing—review and editing; J.V. visualization; J.V. supervision; J.K. (Jiri Koziorek) project administration; J.K. (Jiri Koziorek) funding acquisition. Funding: This research was funded by the Ministry of Education of the Czech Republic (Project No. SP2018/170) and by the European Regional Development Fund in the Research Centre of Advanced Mechatronic Systems project, project number CZ.02.1.01/0.0/0.0/16_019/0000867 within the Operational Programme Research, Development and Education. Acknowledgments: This work was supported by the Student Grant System of VSB Technical University of Ostrava, grant number SP2019/118. This work was supported by the European Regional Development Fund in the Research Centre of Advanced Mechatronic Systems project, project number CZ.02.1.01/0.0/0.0/16_019/0000867 within the Operational Programme Research, Development and Education. The work and the contributions were supported by the project SV4507741/2101, ‘Biomedicínskéinženýrskésystémy XIII’. Conflicts of Interest: The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript, or in the decision to publish the results. References 1. Vanus, J.; Martinek, R.; Kubicek, J.; Penhaker, M.; Nedoma, J.; Fajkus, M. Using the PI processbook software tool to monitor room occupancy in smart home care. In Proceedings of the 2018 IEEE 20th International Conference on e-Health Networking, Applications and Services (Healthcom), Ostrava, Czech Republic, 17–20 September 2018. [CrossRef] 2. Clements-Croome, D.J. Intelligent Buildings; ICE Publishing: London, UK, 2013; p. 1.