Citation: Skålvik, A.M.; Bjørk, R.N.; Martínez, E.; Frøysa, K.-E.; Saetre, C. Multivariate, Automatic Diagnostics Based on Insights into Sensor Technology. J. Mar. Sci. Eng. 2024,12, 2367. https://doi.org/10.3390/ jmse12122367 Academic Editor: Marc Le Menn Received: 31 October 2024 Revised: 17 December 2024 Accepted: 17 December 2024 Published: 23 December 2024 Copyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). Article Multivariate, Automatic Diagnostics Based on Insights into Sensor Technology Astrid Marie Skålvik 1,* , Ranveig N. Bjørk 2, Enoc Martínez 3, Kjell-Eivind Frøysa 4and Camilla Saetre 1 1Department of Physics and Technology, University of Bergen, 5020 Bergen, Norway; camilla.satr[email protected] 2NORCE Norwegian Research Center, 5838 Bergen, Norway;
[email protected] 3SARTI-MAR Research Group, Electronics Department, Universitat Politècnica de Catalunya (UPC), 08800 Vilanova i la Geltrú, Spain; [email protected] 4Department of Computer Science, Electrical Engineering and Mathematical Sciences, Western Norway University of Applied Sciences, 5020 Bergen, Norway; kjell.eivind.fr[email protected] *Correspondence: [email protected] Abstract: With the rapid development of smart sensor technology and the Internet of things, ensuring data accuracy and system reliability is paramount. As the number of sensors increases with demand for high-resolution, high-quality input to decision-making systems, models and digital twins, manual quality control of sensor data is no longer an option. In this paper, we leverage insights into sensor technology, environmental dynamics and the correlation between data from different sensors for automatic diagnostics of a sensor node. We propose a method for combining results of automatic quality control of individual sensors with tests for detecting simultaneous anomalies across sensors. Building on both sensor and application knowledge, we develop a diagnostic logic that can automatically explain and diagnose instead of just labeling the individual sensor data as “good” or “bad”. This approach enables us to provide diagnostics that offer a deeper understanding of the data and their quality and of the health and reliability of the measurement system. Our algorithms are adapted for real time and in situ operation on the sensor node. We demonstrate the diagnostic power of the algorithms on high-resolution measurements of temperature and conductivity from the OBSEA observatory about 50 km south of Barcelona, Spain. Keywords: environmental monitoring; measurement errors; measurement uncertainty; ocean salinity; reliability fault diagnostics; self-validating 1. Introduction Automatic sensor quality control is an integral part of any system of smart sensors. Before using data for updating models, creating predictions or for decision making, data must be checked to ensure they are of high enough quality. For automatic quality control of oceanographic sensors operating underwater, the main approach has been to flag data with quality labels, such as “pass”/“good” data, “not evaluated”, “suspect”/“probably bad”, “fail”/“bad” data and similar ones [ 1 – 3 ]. While the tests for producing these labels are easy to implement, the labels are rather coarse, and it is not always stated in the metadata which thresholds have been used for the different labels. This reduces data re-usability: data that are discarded as they are of too low quality for the original application may be good enough for another, and what is considered noise for some users may be the signal of interest for other users in different domains [ 4 ]. Moreover, the established tests are mainly focused on anomalies in single variables, and even though [ 1 ] suggests combining different variables through a multivariate test, it notes that such testing is challenging and considered experimental. Multivariate tests are not included in publicly available libraries for oceanographic quality control algorithms such as [ 5 – 7 ], and we could not find any documentation that such tests are used in practice elsewhere. J. Mar. Sci. Eng. 2024,12, 2367. https://doi.org/10.3390/jmse12122367 https://www.mdpi.com/journal/jmse
J. Mar. Sci. Eng. 2024,12, 2367 2 of 17 The example provided in [ 1 ] considers detecting a simultaneous high rate of change in temperature and a second variable such as salinity, with thresholds for each variable needed to be set by the operator. However, to our knowledge, there is no documented literature on the application of a multivariate algorithm that not only combines insights into the interrelations between variables, but also addresses the challenge of setting comparable thresholds across different oceanic variables to produce concrete diagnoses. If one extends the scope out of the marine environmental monitoring domain, various methods have been proposed for performing the automatic quality control of sensor systems, ranging from simpler algorithms to sophisticated machine learning approaches for anomaly detection, such as [ 8 – 11 ], and also adapted for multi-sensor applications [ 12 , 13 ]. Common for many of the recent approaches is that they are data-driven, enabled by the rapid increase in computation power combined with access to large datasets. Data-driven quality control can be efficient and powerful for cases where enough data are available. However, if few or no labeled data are available, the data-driven approaches become more complex and less powerful. The complexity increases further for applications with a high correlation between the different data sources, in addition to a high temporal correlation, and when the frequency and duration of anomalies cannot readily be estimated before the system is in operation [ 13 ]. In contrast to process-control systems where operating conditions often are close to training conditions, the conditions in environmental monitoring systems are dynamic, and both the values and statistical properties of the involved parameters may change substantially on different time-scales. Rule-based or model-based diagnostic systems, on the other hand, are based on expert knowledge and are domainand application-dependent [ 14 , 15 ]. Extracting the required information from experts is often a laborious and time-consuming process, though there are several attempts to make this process more effective [ 16 ]. One of the arguments for using data-driven methods for automatic sensor quality control instead of methods relying on expert knowledge is that domain-specific knowledge is not required. As the complexity of the data-driven methods increases, we fall into the other extreme: setting up powerful enough data-driven methods necessitates an expert data scientist. Other important challenges with data-driven methods are that they lack transparency and explainability, although there is ongoing research to overcome this [16]. The lack of oceanographic datasets labeled with diagnosis as opposed to simple “Good”/“Bad” labels, makes supervised machine-learning approaches not applicable as these methods depend on large quantities of labeled data. Unsupervised machine-learning approaches, on the other hand, would identify anomalies but not automatically produce transparent and explainable diagnostics. In this paper, we therefore propose a method for multivariate, rule-based, expertinformed automatic sensor diagnostics, tailored for in situ operation on a sensor node. A sensor node consists of multiple sensors sharing a common prepossessing and communication unit, also referred to as a multiprobe sensor. We provide a framework for transferring domain-expert insights on sensor technology, environmental effects and relation between different environmental and sensed variables into algorithms. With this method, insights into how different environmental conditions and internal errors affect the signals of different sensors, as well as insights into the relation between data recorded by the different sensors, are combined for explainable diagnostics. When an anomaly is detected in a variable measured by one sensor, we use the presence or absence of a similar anomaly in correlated sensor data to automatically propose a physics-informed diagnosis. We demonstrate the method on a sensor node consisting of a temperature, a conductivity and a pressure sensor, deployed underwater. However, the method is applicable to any multi-sensor system with correlated data. 2. Materials and Methods We distinguish between anomaly,event and error. Depending on the thresholds that are used, a detected anomaly is not necessarily a measurement error. In the dynamic
J. Mar. Sci. Eng. 2024,12, 2367 3 of 17 environment encountered in the ocean, especially close to the coast and close to the surface, a statistical outlier in the measurement can in many cases reflect a real change in the water body. Such real changes we here denote as events. If test thresholds are so strict that only nonphysically rapid changes are detected, or if measurement values are outside the possible range, then the anomaly is probably an error from an erroneous measurement. We first detail how anomalies can be detected in the data from each individual sensor (Section 2.1). We propose a combination of algorithms that is robust against missing data, out-of-range values and that can continue to run automatically even after periods with erroneous data. In Section 2.2, we describe different ways for determining whether an anomaly detected in the data from one sensor is also present in data from a sensor measuring a correlated parameter. In Section 2.3, we show how different combinations of anomalies across the variables can be explained by different error modes or by the environment dynamics, and how this information can be used to set up robust diagnostics. The overall method is illustrated in Figure 1. Figure 1. Schematic overview of the proposed method. The method can be applied to any system where the data recorded by different sensors are correlated. In Section 2.4, we describe how the method can be applied to an example measurement system where temperature, conductivity and pressure data are measured, and salinity is calculated based on these measurements. Even though most automatic quality control operations today are performed once the sensors’ data have been transferred to servers, there are a number of advantages in running such operations in situ on the sensor node, before the data are transferred. For sensor nodes relying on wireless communication, in particular for acoustic communication under water, the energy cost of transferring data is high, and much of the data processing needs to be performed at the sensor node level, so that only good data, alarms or diagnoses are communicated onshore. Our algorithms are therefore tailored to work in real-time and in situ, under the following restrictions: • Only the N last points are available for identifying anomalies, where N depends on the smart sensor node’s storing capacity and processing power. • Thresholds cannot be explored and tweaked until the detected anomalies correspond well with what is visually perceived as anomalies. However, if two-way communication is available for the sensors, thresholds can be adjusted after an initial period. The R-program code with algorithms described in this section are available as Supplementary Materials. 2.1. Detection of Anomalies in Individual Variables Depending on the measurement technology and the error condition, different errors can produce different effects on the measurement data, such as drift, increased noise, attenuated measurement signal, flat line, saturation and others [ 17 , 18 ]. Organizations
J. Mar. Sci. Eng. 2024,12, 2367 4 of 17 working with oceanographic measurements, such as [ 1 – 3 ], propose basic algorithms for checking if parameters are outside a range, for spike detection of single value spikes and for detecting high rates of change. We propose an algorithm incorporating elements from these basic tests, with logic for making the algorithm robust against missing data and periods with spiky data. In order to limit ourselves to a few examples for illustration, we therefore here focus on the detection of spikes: where one or several data values are significantly different from the adjacent values, as well as data with a high rate of change: when data values change anomalously much from one timestep to the next. The algorithm is illustrated in Figure 2and is structured as follows: 1. Test if a data point yiis marked as not a number (NaN) or not available (NA). 2. Out-of-range test: check if yi is outside a set range, which could depend on location and depth. 3. High-rate test: if yi and yi−1 have passed both tests 1 and 2, and the difference between yiand yi−1is above a set threshold ydi f f ,max, then yiis marked with “high rate”. 4. Spike test: If a minimum number Nof data points are available, then yiis compared with the moving mean yavg and a multiplier k of the standard deviation σN of the N recent data points that are counted as valid. In order for the algorithm to be robust and not give false detection after periods with very stable conductivity readings and therefore low σN , a minimum variation should be tolerated, here denoted σnat . If one of the following conditions are met, then yiis marked with “spike”: yi<yavg −max⟨k·σN,σnat⟩ yi>yavg +max⟨k·σN,σnat⟩(1) The algorithm relies both on static and dynamic thresholds. For the out-of-range and the high-rate tests, absolute thresholds are set based on application knowledge of what is considered anomalous at the specific location, depth and time of year, taking into account the specific sensor instrumentation. As discussed further in Section 2.2, once an absolute threshold is set for the high-rate test for one variable, corresponding thresholds can be calculated for correlated variables. The dynamic thresholds for spike detection are partly based on the method described by [ 19 ]: a spike is diagnosed based on the running average ¯ yN and a multiple kof the running standard deviation σN , calculated only from the N last data-points that have passed the tests for NA, out-of-range and high-rate. The influence of data points characterized as spikes on the moving mean and standard deviation can be adjusted using an influence parameter w. If wis set to 0, then these data points will not have any effect on the moving statistics. Figure 2. Algorithm for detecting symptoms in individual variables. yi refers to the measurement of variable y at timestep i . N refers to the length of the running window. k is a multiplier used to set dynamic thresholds based on the standard deviation. f ilteredy refers to the N recent data points evaluated as valid and used for calculating statistics such as the mean yavg and the standard deviation σN . ydi f f ,max refers to the absolute threshold set for detecting high rates of change. σnat is the minimum standard deviation that should be accepted, due to natural variations in the environment the sensor is located in.
J. Mar. Sci. Eng. 2024,12, 2367 5 of 17 Note that the use of the statistics of the Nprevious points to evaluate if data point i is a spike or not is vulnerable to any erroneous measurements as the logging begins. The diagnosis algorithm should therefore only be started when the sensor is clean, completely deployed and stabilized into the environment. 2.2. Evaluating If an Anomaly Is Present Across Correlated Variables When taking advantage of multivariate sensor data for diagnostics, one important part of separating between sensor faults or natural events is to determine whether an anomaly is only present in the data from one sensor, or from two or more of the sensors measuring correlated variables. One challenge for detecting if an anomaly is present across different sensors is to set thresholds for detection that are comparable across the variables. In this section, we show how the sensitivity coefficients of the correlated variables can be used to translate a threshold set for one variable into a compatible threshold for a correlated variable. We then explore the use of the co-variation between variables for indicating simultaneous events. Finally, we show how using a running window when comparing variables with slightly different time responses can make the detection of simultaneous signals across variables more robust. 2.2.1. Sensitivity Coefficients for the Correlated Variables In order to set thresholds so that simultaneous anomalies are correctly detected across the different parameters, it is useful to know how sensitive one of the variables are to changes in the other variables. The Guide to the Expression of Uncertainty in Measurement [ 20 ] gives the following definition for sensitivity coefficients for an output estimate ywith respect to an input estimate xi as ci≡δf δxi , where f is the functional relationship for determining the output estimate y . We change the notation slightly here to allow us to distinguish between different output estimates yj , where fj is the functional relationship for determining the output estimate yj:cj,i≡δfj δxi The use of sensitivity coefficients ensures that thresholds can be set consequently across correlated variables. If a threshold x1,max is set for variable 1, the threshold for a correlated variable x2 can be calculated as x2,max =c2,1 ·x1,max . If an anomaly is already detected in variable x1 , and one wants to investigate whether there is a similar anomaly in x2 , a margin could be added so that if the change in variable x2 is slightly attenuated by any other changes in the environment, then the change in x2 is still detected. For example, this can be handled by using the lower 99 % confidence interval bound for the predicted sensitivity coefficient, clwr 1,2 . 2.2.2. Running Co-Variation as an Indicator of Simultaneous Events Both the running correlation coefficient or the running covariance between correlated variables can be used as an indication of overlapping events. It is challenging to set the length of the running window so that dips in the correlation coefficient are discovered timely, without being subject to too much noise. Contrary to setting a threshold for detecting a change in the running correlation, a threshold for detecting a significant covariance can be set based on a threshold chosen for one of the variables, multiplied with a derived threshold for the correlated variable, calculated using the sensitivity coefficients discussed above. 2.2.3. Running Windows for Comparing Signals Across Variables For sensor systems producing high-resolution data, it may be beneficial to add a running window over which the numbers of anomalous data points detected in each variable are counted. This allows for the detection of simultaneous events across variables, even if there are some time delays between the sensor responses. However, if the window lengths are too long, diagnostic information may become lost. Other key factors to consider when selecting window lengths are sensor response time and the environmental dynamics.
J. Mar. Sci. Eng. 2024,12, 2367 6 of 17 2.3. Combining Test Results from Different Correlated Variables to Validate and Set Diagnosis Once algorithms are set up for detecting anomalies in single-sensor data (Section 2.1) and evaluating if anomalies are present across sensors (Section 2.2), the detected anomalies and covariances are automatically combined through a diagnostic logic to set a specific diagnosis. Establishing the diagnostic logic is highly dependent on thorough knowledge of the sensor system. The diagnostic logic for a system with Nx variables, where each can have symptoms ranging from S1 to SNs , with diagnoses D1 to DNd is illustrated in Figure 3 as a flow chart, and in table form in Table 1. Figure 3. Schematic overview of how different symptoms detected in individual variables can be combined with tests for covariance between the variables, into different diagnoses, through a diagnostic logic module. Table 1. Generic diagnostic logic table illustrating symptoms ( Sj ), for different variables Xi , covariances between different variables, and diagnoses ( Dk ). The label “0” indicates no symptoms detected. For the covariance columns, “1” or “0” indicate whether a significant covariance is detected or not, and “-” indicates that it is not relevant for the specific diagnosis. Variable X1X2 ... XNcov(X1,X2) cov(..,..) cov(Xi,XNx) Diagnosis 1 (D1)S1S2 ... - 1 ... 0 S2S1 ... - 1 ... 0 Diagnosis 2 (D2)-S3 ... 0 - ... 1 S1- ... S30 ... 1 ... ... ... ... ... ... ... ... Diagnosis “No Detection” (DNoDet)S1S1 ... ... ... ... 1 S2S2 ... ... ... ... 1 Diagnosis Nd (DNd)... ... ... ... ... ... ... The combination of anomalies from individual variables into a diagnosis will at the same time (a) assign a diagnosis that can help the user deciding whether or not to include the data in his/her application or help the operator determine if any action must be taken for maintaining the sensor node and (b) distinguish between anomalies due to environmental dynamics and due to sensor errors. If two variables X1 and X2 have a strong positive correlation, and a symptom S1 (for example, a positive shift) is observed in X1 , whereas a symptom S2 (for example, a negative shift) is observed in X2 , this would indicate an anomaly due to sensor error, and be assigned a diagnosis D1 . On the other hand, if the same symptom S1 or S2 is observed in both variables, this could indicate a true change in the environment and obtain a “No detection” or similar diagnosis.
J. Mar. Sci. Eng. 2024,12, 2367 7 of 17 2.4. Example Application: CTD Measurement Node For illustration in this paper, we use measured temperature, conductivity and pressure data, as well as calculated salinity data, from the OBSEA observatory in Barcelona, Spain [ 21 ]. The dataset contains data from two Sea-Bird Scientific CTD sensor nodes sequentially deployed, an SBE16 and an SBE37, at a depth of approximately 20 m. When one sensor node is in operation, the other one is onshore for maintenance, and the sensor nodes are never deployed at the same time. As the main differences between the SBE16 and SBE37 sensor nodes are related to battery, data storage and communication, not directly influencing the measurement results when integrated in a cabled observatory, we do not separate between the SBE16 and SBE37 in the presentation of results and in the discussion. Since OBSEA is a cabled observatory, data are streamed to the shore station in real-time, where they are processed and archived. Both sensors have a sampling rate of 10 s. More details are found in [ 21 ], and on the SeaBird website [ 22 , 23 ]. Figure 4shows how the sensor node is installed on the observatory, close to the sea floor. Unprocessed high-frequency data from both sensors can be found at OBSEA’s ERDDAP data service [24]. For illustration, we focus on errors that were expected to occur in conductivity sensors and in the derived salinity. However, this discussion could be extended to include errors in temperature, pressure sensors and others and to include other errors arising from environmental effects or from internal sources such as electronic drift, low battery or other internal malfunctioning. Figure 4. Photo of the SeaBird SBE37SMP sensor node upon installation at the OBSEA observatory. The conductivity sensor studied here has an electrode-based measurement principle. The conductivity 1 ρ is calculated from the measured conductance G and the known cell geometry in terms of length l and cross-sectional area A following the relation derived from [25] (p. 143): 1 G=ρl A Ref. [ 25 ] (p. 145) notes that while “steady, biological growth” may result in a linear drift in the conductivity, “biological settling or an increase in biological productivity” can result in a more episodic change. The measurement technology of the conductivity sensor studied here consists of a protected conductivity cell into which water is pumped and sampled at set intervals. The conductivity measurement is therefore approximate to both temporal and spatial averages
J. Mar. Sci. Eng. 2024,12, 2367 8 of 17 of the conductivity of the seawater in the proximity of the pump. Any rapid, instantaneous or local changes in conductivity is smoothed out by that measurement process. Salinity was calculated from the measured temperature, conductivity and pressure. Refs. [ 25 , 26 ] (p. 145) point out that an offset between the measured temperature and the actual temperature in the conductivity cell may lead to spikes in the calculated salinities. 2.4.1. Detection of Anomalies in Individual Variables Tests for anomalies in individual variables were carried out for both measured and derived parameters. Salinity was a derived parameter, where errors could stem from both the measured variables used to calculate it (such as temperature, conductivity and pressure), but also from the calculation process itself, where for example asynchronous input data may be an issue. The threshold for the rate of change in temperature was set to 0.05 ◦ C and translated into corresponding thresholds for conductivity and salinity using the calculated sensitivity coefficients as detailed in the following section. Dynamic thresholds for spike detection were calculated dynamically as described in Section 2.1, with k= 4 times σN as the number of standard deviations outside the average, but allowing for a minimum of natural variations. 2.4.2. Evaluating if an Anomaly Is Present Across Correlated Variables Sensitivity coefficients: The UNESCO [ 27 ] (pp. 6–12) equations of state allow the calculation of practical salinity as a function of conductivity, temperature and pressure, and conductivity as a function of salinity, temperature and pressure. We calculated the local sensitivity coefficients according to the UNESCO equations of state implemented in the R-package “oce” [ 28 ] for the conductivity and salinity calculations (functions “swSCTp” and “swCSTp”, respectively). Figure 5shows the sensitivity coefficients calculated for salinity and conductivity with respect to temperature. As a practical approach in the diagnostic algorithms, in the examples shown in this paper, we estimated the sensitivity coefficients based on linear and quadratic regression models built on a subset of recorded data from 2020, excluding obvious outliers. When there was an increase in temperature of 1 ◦ C from approximately 20 ◦C , but the conductivity measurement was constant, this resulted in a reduction in the calculated salinity of approximately − 0.9 PSU (Figure 5a). An increase in temperature of 1◦C , at constant salinity, resulted in an increase in the measured conductivity of approximately 0.1 S/m (Figure 5b).
J. Mar. Sci. Eng. 2024,12, 2367 9 of 17 Figure 5. Absolute sensitivity coefficients for (a) salinity and (b) conductivity with respect to temperature. Calculated for every 1000th data point (approximately every 3 h) in the OBSEA CTD data for year 2020. The dashed lines indicate the 99 percent confidence levels of the predicted sensitivity coefficients using a quadratic (a) and linear (b) model built on the calculated sensitivity coefficients, excluding extreme outliers. Running windows: For the 10 s resolution system studied in this paper, we chose a window length of 30 min ( N= 30 ·60 10 ) for the running statistics discussed in Section 2.1 . The covariance over different window lengths is shown in Figure 6. For comparing symptoms across variables, we used a running window of 4 data points. Figure 6. Rolling co-variation between temperature and conductivity over 15 min, 1 h and 2 h, for 5–6 August 2020. 2.4.3. Combining Tests Results for Setting Diagnosis Based on the insights into possible error sources and effects on the involved sensor signals described above, we systematically mapped out combinations of anomalies with
J. Mar. Sci. Eng. 2024,12, 2367 16 of 17 simple Good/Bad flags predominant in the domain today. We believe this can be valuable information for both system operators and data users. Supplementary Materials: The following supporting information can be downloaded at: https://www. mdpi.com/article/10.3390/jmse12122367/s1. R code for multivariate, automatic diagnostics based on insights in sensor technology; data and results for 2020; zip file with weekly plots for 2020. Author Contributions: Conceptualization, C.S. and K.-E.F.; methodology, A.M.S.; software, A.M.S.; formal analysis, A.M.S.; investigation, A.M.S.; data curation, E.M.; writing—original draft preparation, A.M.S.; writing—review and editing, A.M.S., C.S., R.N.B., K.-E.F. and E.M.; visualization, A.M.S.; supervision, C.S., R.N.B. and K.-E.F.; project administration, C.S. and K.-E.F.; funding acquisition, C.S. All authors have read and agreed to the published version of the manuscript. Funding: This work is part of the SFI Smart Ocean (a Centre for Research-based Innovation). The Centre is funded by the partners in the Centre and the Research Council of Norway (project no. 309612). Data Availability Statement: The original data presented in the study are openly available in the OBSEA data repository, at https://data.obsea.es/erddap/tabledap/OBSEA_CTD_full.html (accessed on 10 December 2024). Conflicts of Interest: The authors declare no conflicts of interest. References 1. IOOS and QARTOD. Manual for Real-Time Quality Control of In-Situ Temperature and Salinity Data. 2020. Available online: https://ioos.noaa.gov/ioos-in-action/temperature-salinity/ (accessed on 2 April 2024) . 2. Wong, A.; Keeley, R.; Carval, T. Argo Quality Control Manual for CTD and Trajectory Data; Report; Argo Data Management Team: Brest, France , 2024. [CrossRef] 3. EuroGOOS DATA-MEQ Working Group. Recommendations for In-Situ Data Near Real Time Quality Control. Report , 2010. Available online: https://archimer.ifremer.fr/doc/00251/36230/ (accessed on 24 October 2024). 4. Nguyen, N.T.; Lima, K.; Skålvik, A.M.; Heldal, R.; Knauss, E.; Oyetoyan, T.D.; Pelliccione, P.; Sætre, C. Synthesized Data Quality Requirements and Roadmap for Improving Reusability of In-Situ Marine Data. In Proceedings of the 2023 IEEE 31st International Requirements Engineering Conference (RE), Hannover, Germany, 4–8 September 2023; pp. 65–76. [CrossRef] 5. IOOS QC: QARTOD and Other Quality Control Tests Implemented in Python. 2022. Available online: https://pypi.org/project/ ioos-qc/ (accessed on 24 October 2024). 6. Python Functions defined for computational ION. 2017. Available online: https://github.com/ooici/ion-functions/tree/master/ ion_functions/qc (accessed on 24 October 2024). 7. Lookup Tables for the Automated OOI Quality Control Algorithms. 2024. Available online: https://github.com/oceanobservatories/ qc-lookup (accessed on 24 October 2024). 8. Barbariol, T.; Chiara, F.D.; Marcato, D.; Susto, G.A. A Review of Tree-Based Approaches for Anomaly Detection. In Control Charts and Machine Learning for Anomaly Detection in Manufacturing; Tran, K.P., Ed.; Springer International Publishing: Cham, Switherland, 2022; pp. 149–185. [CrossRef] 9. Teh, H.Y.; Wang, K.I.K.; Kempa-Liehr, A.W. Expect the Unexpected: Unsupervised Feature Selection for Automated Sensor Anomaly Detection. IEEE Sens. J. 2021,21, 18033–18046. [CrossRef] 10. Han, X.; Jiang, J.; Xu, A.; Bari, A.; Pei, C.; Sun, Y. Sensor Drift Detection Based on Discrete Wavelet Transform and Grey Models. IEEE Access 2020,8, 204389–204399. [CrossRef] 11. Zhu, M.; Li, J.; Wang, W.; Chen, D. Self-Detection and Self-Diagnosis Methods for Sensors in Intelligent Integrated Sensing System. IEEE Sens. J. 2021,21, 19247–19254. [CrossRef] 12. Yan, X.; Yan, W.J.; Xu, Y.; Yuen, K.V. Machinery multi-sensor fault diagnosis based on adaptive multivariate feature mode decomposition and multi-attention fusion residual convolutional neural network. Mech. Syst. Signal Process. 2023,202, 110664. [CrossRef] 13. Belay, M.A.; Blakseth, S.S.; Rasheed, A.; Salvo Rossi, P. Unsupervised Anomaly Detection for IoT-Based Multivariate Time Series: Existing Solutions, Performance Analysis and Future Directions. Sensors 2023,23, 2844. [CrossRef] [PubMed] 14. Angeli, C. Diagnostic Expert Systems: From Expert’ s Knowledge to Real-Time Systems. In Advanced Knowledge Based Systems: Model, Applications & Research; TMRF e-Book; TMRF: Burntwood, UK, 2010; Volume 1, pp. 50–73. 15. Gao, Z.; Cecati, C.; Ding, S.X. A Survey of Fault Diagnosis and Fault-Tolerant Techniques—Part II: Fault Diagnosis With Knowledge-Based and Hybrid/Active Approaches. IEEE Trans. Ind. Electron. 2015,62, 3768–3774. [CrossRef] 16. Young, A.; West, G.; Brown, B.; Stephen, B.; Duncan, A.; Michie, C.; McArthur, S.D. Parameterisation of domain knowledge for rapid and iterative prototyping of knowledge-based systems. Expert Syst. Appl. 2022,208, 118169. [CrossRef]
J. Mar. Sci. Eng. 2024,12, 2367 17 of 17 17. Skålvik, A.M.; Saetre, C.; Frøysa, K.E.; Bjørk, R.N.; Tengberg, A. Challenges, limitations, and measurement strategies to ensure data quality in deep-sea sensors. Front. Mar. Sci. 2023,10, 1152236. [CrossRef] 18. Jesus, G.; Casimiro, A.; Oliveira, A. A Survey on Data Quality for Dependable Monitoring in Wireless Sensor Networks. Sensors 2017,17, 2010. [CrossRef] [PubMed] 19. Brakel, J. Robust Peak Detection Algorithm Using Z-Scores. 2020. Available online: https://stackoverflow.com/questions/2258 3391/peak-signal-detection-in-realtime-timeseries-data/2264036222640362 (accessed on 15 January 2024). 20. BIPM; IEC; IFCC; ILAC; ISO; IUPAC; IUPAP; OIML. Evaluation of Measurement Data—Guide to the Expression of Uncertainty in Measurement. Joint Committee for Guides in Metrology, JCGM 100:2008. Available online: https://www.bipm.org/documents/ 20126/2071204/JCGM_100_2008_E.pdf/cb0ef43f-baa5-11cf-3f85-4dcd86f77bd6 (accessed on 12 November 2023). 21. Del-Rio, J.; Nogueras, M.; Toma, D.M.; Martínez, E.; Artero-Delgado, C.; Bghiel, I.; Martinez, M.; Cadena, J.; Garcia-Benadi, A.; Sarria, D.; et al. Obsea: A Decadal Balance for a Cabled Observatory Deployment. IEEE Access 2020,8, 33163–33177. [CrossRef] 22. Scientific, S. SBE 16plus V2 SeaCAT. Available online: https://www.seabird.com/sbe-16plus-v2-seacat/product?id=60761421598 (accessed on 9 December 2024). 23. Scientific, S. SBE 37-SM, SMP, SMP-ODO MicroCAT. Available online: https://www.seabird.com/moored/sbe-37-sm-smp-smpodo-microcat/family?productCategoryId=54627473786 (accessed on 9 December 2024). 24. Martínez, E. OBSEA ERDDAP Data Service. 2024. Available online: https://data.obsea.es/erddap (accessed on 10 December 2023). 25. Venkatesan, R.; Tandon, A.; D’Asaro, E.; Atmanand, M. Observing the Oceans in Real Time; Springer: Cham, Switzerland, 2018; pp. 144–145. 26. Jansen, P.; Weeding, B.; Shadwick, E.H.; Trull, T.W. IMOS—Southern Ocean Time Series (SOTS)—Quality Assessment and Control Report Temperature Records; Technical Report; Commonwealth Scientific and Industrial Research Organisation (CSIRO): Tasmania, Australia, 2021. [CrossRef] 27. Fofonoff, N.; Millard, R., Jr. Algorithms for computation of fundamental properties of seawater. Unesco Tech. Pap. Mar. Sci. 1983, 44 . [CrossRef] 28. Kelley, D.E. Oceanographic Analysis with R; Springer: New York, NY, USA, 2018. [CrossRef] 29. Sea-Bird Scientific. Sea-Bird Scientific University Module 12: Advanced Data Processing: Dynamic Corrections for CTDs. Available online: https://www.seabird.com/training-materials-download (accessed on 10 October 2024). 30. Komadina, A.; Martini´c, M.; Groš, S.; Mihajlovi´c, Ž. Comparing Threshold Selection Methods for Network Anomaly Detection. IEEE Access 2024,12, 124943–124973. [CrossRef] 31. Nguyen, N.T.; Skalvik, A.M.; Sylligardos, E.; Heldal, R.; Pelliccione, P.; Boniol, P.; Palpanas, T.; Alvsvag, S. Interpretable Multivariate Anomaly Detector Selection for Automatic Marine Data Quality Control. In Proceedings of the 41st IEEE International Conference on Data Engineering—Industrial Track (ICDEIndustrial2025), Hong Kong, 19–23 May 2025; Submitted. 32. Zang, C.; Huang, S.; Wu, M.; Du, S.; Scholz, M.; Gao, F.; Lin, C.; Guo, Y.; Dong, Y. Comparison of Relationships Between pH, Dissolved Oxygen and Chlorophyll a for Aquaculture and Non-aquaculture Waters. Water Air Soil Pollut. 2011,219, 157–174. [CrossRef] Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.