scieee AI-readable full text Open interactive document viewer

M2.4/3.3 v2 Submission of developed models for publication

Varouchakis, Emmanouil

Abstract

The research project is implemented in the framework of H.F.R.I call “Basic research Financing (Horizontal support of all Sciences)” under the National Recovery and Resilience Plan “Greece 2.0” funded by the European Union – NextGenerationEU (H.F.R.I. Project Number: 16537).

Full text

Mathematical Geosciences Enhancing Geostatistical Analysis of NaturalResources Data with Complex Spatial Formations through non-Euclidean Distances --Manuscript Draft-- Manuscript Number: MATG-D-24-00163 Full Title: Enhancing Geostatistical Analysis of NaturalResources Data with Complex Spatial Formations through non-Euclidean Distances Article Type: S.I. : GeoENV2024 Keywords: geostatistics; covariance models; non-Euclidean distance; MLE; Gaussian anamorhposis Corresponding Author: Emmanouil Varouchakis Technical University of Crete: Polytechneio Kretes CHANIA, GREECE Corresponding Author Secondary Information: Corresponding Author's Institution: Technical University of Crete: Polytechneio Kretes Corresponding Author's Secondary Institution: First Author: Maria Despoina Koltsidopoulou, M.Sc. First Author Secondary Information: Order of Authors: Maria Despoina Koltsidopoulou, M.Sc. ANDREAS PAVLIDES, Ph.D Dionissios T. Hristopulos, PhD Emmanouil A. Varouchakis, PhD Order of Authors Secondary Information: Funding Information: Hellenic Foundation for Research and Innovation (16537) Prof. Emmanouil A. Varouchakis Abstract: In geostatistical modeling, traditional covariance models based on Euclidean distance often fail in complex geographic environments, especially when natural barriers or irregular terrain distort spatial relationships. This study explores the use of nonEuclidean distance metrics, such as Manhattan, Minkowski, and Chebyshev distances, to improve the accuracy of spatial covariance models in such environments. We apply several covariance models -Exponential, Gaussian, Spherical, Spartan Spatial Random Field (SSRF), and the recently developed Harmonic Covariance Estimation (HCE) modelto estimate soil aluminium concentration data from a region with complex physical barriers. Gaussian Anamorphosis with the new Kernel Cumulative Density Estimator (KCDE) was used to normalize the data. The model parameters for all covariance models and all distances were estimated via maximum likelihood estimation (MLE). The models were tested for their applicability in non-Euclidean distances by calculating the Eigenvalues of the resulting covariance matrices, in order to ensure positivedefiniteness. In this case study, the Gaussian and Spherical Models were found to be inadequate for non-Euclidean spaces as the covariance matrices were not positive-definite. On the other hand, the SSRF, Exponential, and HCE models provided valid covariance matrices consistently. In particular, the HCE model outperformed the others in all distance metricsdemonstrating the robustness and adaptability of HCE in a nonEuclidean geostatistical environment. Powered by Editorial Manager® and ProduXion Manager® from Aries Systems Corporation Enhancing Geostatistical Analysis of Natural Resources Data with Complex Spatial Formations through non-Euclidean Distances Maria Despoina Koltsidopoulou1, Andrew Pavlides2, Dionissios T. Hristopulos1, Emmanouil A. Varouchakis2* 1School of Electrical and Computer Engineering, Technical University of Crete, Chania, 73100, Crete, Greece. 2School of Mineral Resources Engineering, Technical University of Crete, Chania, 73100, Crete, Greece. *Corresponding author. E-mail: evarouc[email protected]; Abstract In geostatistical modeling, traditional covariance models based on Euclidean distance often fail in complex geographic environments, especially when natural barriers or irregular terrain distort spatial relationships. This study explores the use of non-Euclidean distance metrics, such as Manhattan, Minkowski, and Chebyshev distances, to improve the accuracy of spatial covariance models in such environments. We apply several covariance models -Exponential, Gaussian, Spherical, Spartan Spatial Random Field (SSRF), and the recently developed Harmonic Covariance Estimation (HCE) modelto estimate soil aluminium concentration data from a region with complex physical barriers. Gaussian Anamorphosis with the new Kernel Cumulative Density Estimator (KCDE) was used to normalize the data. The model parameters for all covariance models and all distances were estimated via maximum likelihood estimation (MLE). The models were tested for their applicability in non-Euclidean distances by calculating the Eigenvalues of the resulting covariance matrices, in order to ensure positive-definiteness. In this case study, the Gaussian and Spherical Models were found to be inadequate for non-Euclidean spaces as the covariance matrices were not positive-definite. On the other hand, the SSRF, Exponential, and HCE models provided valid covariance matrices consistently. In particular, the HCE model outperformed the others in all distance metrics demonstrating the robustness and adaptability of HCE in a non-Euclidean geostatistical environment. 1 Manuscript Click here to access/download;Manuscript;main.tex Click here to view linked References 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 Keywords: geostatistics, covariance models, non-Euclidean distance, MLE, Gaussian anamorhposis 1 Introduction Traditional covariance models based on Euclidean metrics are likely to fail in complex geographical environments. For instance, spatial relationships are distorted by physical barriers such as faults through which the deposition of natural resources can be disturbed. Standard models based on non-Euclidean distance metrics usually lead to inaccuracies of resources and the reason is these disturbances. Addressing these challenges can occur by incorporating non-Euclidean distance metrics, such as Manhattan or Chebyshev distances, through which more robust and flexible options are provided for analyzing complex geographic domains. This approach therefore increases the possibility of more accurate modelling and improved resource management. However, while traditional covariance models are based on Euclidean distances and are designed to produce positive-definite covariance matrices, they can fail when applied to non-Euclidean spaces. Consequently, we are led to inaccurate statistical analysis and prediction as positive definiteness leads to the invertibility of a covariance matrix. To address such challenges, the development of new covariance models that can be adapted to non-Euclidean measures arises. A necessary condition is that the covariance matrix remains positively definite. The study of traditional geostatistical models based on Euclidean distances contains some limitations; in his paper, Frank C. Curriero (2006) advocates using non-Euclidean metrics in cases where physical barriers distort spatial relationships. After analysis and simulations, he concludes that non-Euclidean distances can significantly improve prediction accuracy in complex geographic environments such as river networks and irregular terrains. In particular [1], he explains how eigenvalues can be used to test whether a covariance matrix with non-Euclidean distances is positivedefinite, showing that the exponential model leads to permissible matrices for certain non-Euclidean distances. This approach is further documented in the research conducted by Theodoridou et al. (2017), where the application of Euclidean distance metrics, such as that of Manhattan, was shown to improve the exact prediction in spatial geostatistical analyses [2]. Building on this, the study by BJK Davis et al. (2019) presents a new method to enhance the validity and accuracy of spatial covariance structures based initially on non-Euclidean metrics in geostatistical models. Through simulations and real-world applications, such as water quality assessments in the Chesapeake Bay, this method improves geostatistical predictions [3]. Further research has been carried out by Hristopulos (2023) who introduced advanced methodologies to develop non-separable covariance kernels in Gaussian spatio-temporal processes using a hybrid spectral method. The method is inspired by physical systems as well as the linear damped harmonic oscillator (LDHO) model. More specifically, the model is one that simulates the back and forth motion of a 2 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 system (such as a spring or a pendulum) where a restoring force is compensated by a damping force (such as friction or air resistance). This approach allows kernels to incorporate monotonic and oscillatory damping into their spacetime correlations. The use of the stochastic LDHO model to construct spatio-temporal covariance kernels improves existing models leading to higher prediction accuracy for data governed by non-Euclidean processes. These developments highlight the potential of non-Euclidean approaches to better identify spatial relationships in data in complex systems [4]. Further studies have explored in depth the application of non-Euclidean metrics in geostatistical modeling to address complex spatio-temporal relationships, such as Agou et al. (2019) who demonstrated the use of kriging regression on the island of Crete to analyze rainfall data from a sparse network of monitoring stations. In their research, they were led to anisotropic spatial correlations through the integration of topographic data and Spartan variogram function [5]. On similar lines, Varouchakis et al. (2022) used spatiotemporal geostatistical analysis to quantify groundwater level variations, demonstrating that incorporating the Manhattan distance metric led to a significant improvement in model accuracy in regions with complex geological features [6]. In addition to these, Ver Hoef et al. (2018) developed a reduced-degree predictive process model that constructs valid spatial covariance tables. In the process, they addressed the challenges of implementing kriging models with non-Euclidean distances [7]. The main objective of this study is to investigate existing models and develop a new covariance model that can be reliably used with non-Euclidean distance metrics to improve spatial analysis and resource estimation in complex geographic environments. Focusing on Manhattan and different implementations of Minkowski and Chebyshev distances, our work will explore how these metrics can be incorporated into geostatistical models to provide more accurate and realistic representations of spatial relationships and a comparison of the different models tested. The scope includes the application of the methods to real data of aluminium soil concentration originally taken for a geochemical analysis. Gaussian anamorphosis was deemed necessary. This anamorphosis was performed with a novel method based on the Cumulative Density Function, not the Probability Density Function, and is called Kernel Cumulative Density Estimator (KCDE). 2 Methodology The spatial distribution of variables is used utilizing the mathematical framework of random fields [8–10]. A scalar random field X(s) where s∈Rdrepresents the position vector, is defined as a random function D ⊂ Rd→R, where Dis the spatial domain. This random field is characterized by n-point joint probability density functions, where n= 1,2, . . . denotes a positive integer. The data, represented by the sampling regions {s1,...,sN}and the corresponding values of the random variables {x1, . . . , xN}are considered as partial samples of the random field. The field is reconstructed into a rectangular map grid Gconsisting of points {z1,...,zP}. The estimation points will be denoted in general by u∈Rd, with estimation points u∈G. The random fields are decomposed into a deterministic part (trend) and a stochastic part (fluctuation) as follows: 3 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 X(s) = mX(s) + X′(s) (1) The function mX(s) is the trend. It represents the mathematical expectation of the random field, E[X(s)] and it is often modeled as a low-degree polynomial. The fluctuation X′(s) represents the stochastic component of X(s). It is obtained by subtracting the trend from X(s) and thus is equal to the residuals of the field [11]. 2.1 Distance Measures In this work, various distances are examined and explored along with the covariance models. Euclidean distance: Euclidean distance is the most commonly used form of distance and represents the actual distance in our world. It calculates the straight-line distance between two points P= (p1, p2, ..., pn) and Q= (q1, q2, ..., qn) in an n-dimensional space using the formula: dEuclidean(P,Q) = v u u t n X i=1 (pi−qi)2,(2) Manhattan Distance: Manhattan distance (also known as taxicab or city block), measures the distance between two points by summing the absolute differences in their coordinates, as shown in: dManhattan(P,Q) = n X i=1 |pi−qi|(3) This distance measure is most applicable in cases where traffic or paths are confined to a grid. Typically this is the case with fluids passing through a porous material or water in rivers. Minkowski distance: The Minkowski distance provides a generalized framework for measuring distance in multidimensional spaces via the parameter k. This distance can express both the Euclidean distance (for k= 2) and Manhattan distance (for k= 1) as special cases. The Minkowski distance between points Pand Qis defined as: dMinkowski(P,Q) = n X i=1 |pi−qi|k!1/k .(4) Chebyshev distance: The Chebyshev distance (also known as chessboard distance) calculates the maximum difference between the coordinates of two points along any dimension. The equation 4 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 for the Chebyshev distance between points Pand Qis expressed as follows: dChebyshev(P,Q) = max i(|pi−qi|) (5) This distance measure is employed in scenarios where only the highest difference in any dimension matters, such as in specific optimization problems or game strategies. In particular, Fig. 1presents an illustrative example of the two-dimensional distance for three different distance measures, between the point P= (0,0) (origin) and points Qi= (qi,1, qi,2), i = 1, ..., 9. (a) Euclidean (b) Chebyshev (c) Manhattan Fig. 1: An example of the various distances in a two-dimensional space between point P= (0,0) and points Qi= (qi;1, qi;2)=({0,1,2,0,1,2,0,1,2},{0,0,0,1,1,1,2,2,2}). 2.2 Gaussian Anamorphosis Gaussianity is desirable in the estimation of random fields. Various studies have shown that linear predictors perform better if the underlying distribution is close to the normal distribution [11–13]. This research uses a kernel-based (non-parametric) data-driven approach to conduct Gaussian anamorphosis [12,14], KCDE. Unlike other non-parametric methods that focus on probability density, this method directly targets the Culmulative Density Function (CDF) with explicit density integration. This method has been described in detail [15], but a summary of the process is provided here. Let {x[i]}N i=1 represent the ordered sample of data values. Then, the nonparametric, kernel-based estimate of the CDF (KCDE), ˆ FK(·), can be obtained from the following weighted sum [12]: ˆ FK(x) = N X i=1 1 N˜ Kx−x[i] b(6) 5 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 where ˜ K(·) corresponds to CDF kernel step which is defined via the integral [10,16]: ˜ Kx−x[i] b=1 bZx −∞ Kx′−x[i] bdx′,(7) where Kis the chosen Kernel function, where in this case study is the triweight kernel [16]. Consequently, in the following implementation, the Gaussian anamorphosis is performed on the fluctuation X′(s) (i.e. after the trend is removed) using KCDE in combination with the normal-effects transformation method [17], as explained in [15]. In addition, a look-up table was used in combination with the normal score transformation and KCDE to perform the anamorphosis. 2.3 Covariance Functions The models attempting to represent spatial processes rely on the spatial correlations between the data. These correlations are reflected in the covariance function that characterizes the random field: cX(s1,s2) = E[X(s1)X(s2)] −E[X(s1)] E[X(s2)] (8) The covariance function is fundamental in exploring the spatial or spatiotemporal relationships within the values of the investigated random field, as it quantifies how values at different locations co-vary [18]. However, not all functions can be used as a covariance function. For a function to be permissible as a covariance model, it should satisfy specific mathematical conditions. For Euclidean distances, adherence to the conditions set by the Bochner Theorem [19] ensures that the function is a valid covariance function and thus, it will result in a positive-definite covariance matrix. When dealing with non-Euclidean distances the more generic Cram´er extends the applicability of covariance functions to non-Euclidean spaces and thus, consistent with the underlying geometry of such spaces [20,21]. Commonly used theoretical covariance models include the exponential, Gaussian, and spherical models [22]. The equations for these models are provided below, where σ2 !Xrepresents the variance, |r|denotes the Euclidean norm of the lag vector r, and ξ is the correlation length. Exponential cX(r) = σ2 X[exp (−∥r∥/ξ)] .(9) Gaussian cX(r) = σ2 Xexp −∥r∥2/ξ2.(10) Spherical cX(r) =    σ2 X1−1.5∥r∥ ξ+ 0.5∥r∥ ξ3if ∥r∥ ≤ ξ 0 if ∥r∥ ≥ ξ. (11) 6 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 2.3.1 Spartan spatial random field model Spartan Spatial Random Fields (SSRFs) are a family of geostatistical models [23]. The case of covariance SSRFs is provided, where a Spartan-specific probability density (PDF) with spatial dependence derived from generalized gradient and Laplacian operators is used. The PDF of the SSRF depends on a small set of parameters, allowing modeling of spatial processes with controlled roughness and differentiability [24]. In three spatial dimensions, the SSRF model is given by [25,26]: C(r) =            η0 2π∆E−|r| ξβ2sin( |r| ξβ1) |r| ξ|η1|<2, η0 8πE−|r| ξη1= 2, η0 4π∆E−|r| ξω1−E−|r| ξω21 (ω1−ω2)|r| ξ η1>2. (12) In equation (12), η0is the scale factor that determines the magnitude of the fluctuations, η1is the rigidity coefficient that controls the shape of the covariance function, and ξis the characteristic length that determines the range of spatial correlations and normalizes the distance. The additional coefficients are defined as β1,2=|2∓η1|1/2, ω1,2=η1∓∆ 21/2, and ∆ = |η2 1−4|1/2. The correlation length in the SSRF model is influenced by both η1and ξ. Notably, when η1= 2, the Spartan covariance simplifies to the exponential model [25]. SSRF models are differentiable, a property that is useful for applications where smoothness of the spatial field is important. The SSRF model has performed well in various interpolation studies, such as those involving radioactivity dose rates [23], coal reserves estimation [11], and groundwater level modeling [27]. Furthermore, it has been proven to always give positive-definite covariance matrices for the Manhattan distance [6]. 2.3.2 Covariance in non-Euclidean distances An investigation of the validity of covariance models using non-Euclidean distance metrics was carried out through the study of Curriero et al. (2006) and internally developed code (MATLAB 23a) [1]. This investigation aimed to clarify how different spatial metrics, specifically Manhattan distance, affect the validity and performance of different covariance models. The findings indicated that, unlike Gaussian, spherical, and various forms of the Mat´ern class, the exponential covariance function returns admissible, positively definite covariance matrices with Manhattan distance. Further analysis by [6] revealed that the Spartan Spatial Random Field (SSRF) model, which shares structural similarities with the exponential covariance function, also produces valid spatiotemporal covariance structures when the Manhattan distance is used [25,26]. Therefore, the SSRF covariance has guaranteed consistency for the Manhattan distance. Consequently, the Spartan model is acceptable at certain nonEuclidean distances, which conclusion was supported in a further analysis through the application of the assumption of Curriero et al. (2006) [1]. The results confirmed that the eigenvalues of the covariance matrix derived from the Spartan model are positive, thus ensuring that the covariance matrix remains positively definite. 7 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 Covariance models, such as those of spherical and Gaussian models, do not always produce permissible, positive-definite covariance matrices for non-Euclidean distances. In certain cases, covariances may result in a permissible matrix, which may not be sufficient. There have been several examples in the bibliography where these models could not form a positive definite matrix [1,3,6]. 2.4 HCE: Covariance model based on LDHO In [28], the Linearly Damped Harmonic Oscillator (LDHO) leads to the following spatiotemporal covariance function: C(r, τ) = σ2 0 (2π)d/2rνZ∞ 0 kd/2Jν(kr)A(k)e−|τ|B(k)/τcdk, (13) where rrepresents the spatial distance and τdenotes the temporal lag. The parameter σ2 0is the variance of the process, and dindicates the spatial dimension. The term νis the order of the Bessel function Jν(kr), with Jν(kr) being the Bessel function of the first kind of order ν=d/2−1. The functions A(k) and B(k) define the dependencies on spatial frequency, and τcis the characteristic timescale governing the temporal decay. Different functional forms of A(k) and B(k) yield permissible covariance functions under the Bochner Theorem, which produce positive definite matrices under Euclidean distance. To adapt the LDHO model to non-Euclidean distances, we consider an alternative parametrization where A(k) = exp(−βk) and B(k) = α+ξk with α, β being model parameters and ξthe correlation length. Substituting these into Eq. 13, the covariance function turns into: C(r, τ) = σ2 0e−a|τ|/τc (2π)d/2rνZ∞ 0 kν+1Jν(kr)e−k(β+ξ|τ|/τc)dk. (14) After performing the integration, the resulting spatiotemporal covariance function is: C(r, τ) = σ2 0Γd+1 2 πd+1 2˜τc (β˜τc+ξ|τ|)e−a|τ|/˜τc"r2+β+ξ|τ| ˜τc2#−d+1 2 .(15) For the spatial marginal covariance, assuming isotropy, we obtain the covariance function (referred to as Harmonic Covariance Estimation, HCE): CHCE(r) = σ2 0Γd+1 2β πd+1 2(r2+β2)d+1 2 (16) . This HCE covariance function is permissible in Euclidean space. We applied Curriero’s et al. (2006) [1] method to test the eigenvalues of the covariance matrices derived from under non-Euclidean distances. The matrices remained positive-definite, ensuring their invertibility and applicability in kriging predictions. In Fig. 2, we present a visual comparison of three key covariance functions: the HCE from Eq. (16), the Spherical model, and the Exponential model. This comparison 8 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 It is not uncommon for the back-transform of predictions to include bias [15,17]. This is also the case in our study, with the bias (ME) being significant compared to the range of the original data. One reason for the bias is that when the symmetric predictions are back transformed to the original, skewed distribution, the kriging predictors are aligned with the median of the prediction’s distribution rather than the forecast distribution’s mean. Another factor that increases the mean error is that the highest value (76280 ppm) is significantly higher than the Al concentration at the surrounding sampling locations. Thus, the LCV that works by subtracting a point and predicting the value at that location is disadvantaged when trying to predict the value at that location, as the actual value is an outlier. Some example Kriging maps for HCE and Exponential covariance are shown in Fig. 5. The Kriging search neighborhood depends on the estimated correlation range (i.e. range of influence) and is therefore different for each distance and covariance model. It is shown in Fig. 5that all displayed covariance models and distance metrics effectively capture the spatial variability of the field. All figures indicate the high concentration at the northwestern edge with a small area of high concentration about 3 km east and 4 km north. Low values are captured in the northeastern part in all cases. The similarities between the kriging maps are expected, considering that both Exponential and HCE yielded similarly well-performed LCV results. 5 Conclusions In this study, we tested various covariance models; the Gaussian, Exponential, Spherical, SSRF and the recently developed HCE model with non-Euclidean distance metrics for aluminium concentration in soil data. Sampling was informal and focused on a few easily accessible areas, effectively removing much of the field without sampling. Data were normalized using KCDE Gaussian anamorphosis with a look-up table. The normal probability plot revealed that the transformed data were closely matched to the normal distribution. Using the normalized data, MLE was employed to estimate the parameters of the different tested models for various distance metrics. The metrics tested were Euclidean, Manhattan distance, Chebyshev distance and the Minkowski distance with k= 2/3, k = 1.5. The eigenvalues of the resulting covariance matrices were calculated. The Gaussian and Spherical models failed to produce positive-definite covariance matrices for non-Gaussian distance metrics and were therefore omitted. The SSRF, Exponential and HCE covariance matrices were positive-definite all distances investigated and could therefore be used for kriging predictions. We used these covariance models to perform OK on a grid with a cell size of 100m x 100m. While all three models were able to provide predictions, the validation measures were better for HCE, satisfactory for Exponential, and SSRF underperformed with the exception of Manhattan distance. For both the HCE and the Exponential model, the various distances returned similar validation measures. HCE had RMSEs from 5609 ppm to 5613 ppm with ρ= 0.758 to ρ= 0.759. Validation measures for the Exponential covariance also had a small range with RMSEs from 6243 ppm to 6261 15 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 ppm with ρ= 0.687 to ρ= 0.689. Thus, while using non-Euclidean distances did not significantly improve the predictions in this case, HCE performed well and the non-Euclidean distance metrics were competitive. The bias in the back-transformation for all distances was minor compared to the range of data values (0-76280 ppm), but noticeable (in the range of −540 ppm to −605 ppm). The reason for the bias is the way back-transformation works on skewed random variables and also because LCV is disadvantaged by the presence of outliers with very high values compared to the variance of the field and the value of their neighbors. Given the challenges presented by the informal sampling grid and data distribution, HCE combined with KCDE Gaussian anamorphosis can be considered to give satisfactory and promising results, encouraging further study. In the future, multivariate analysis using co-kriging could be employed and different datasets could be investigated to explore the robustness of the HCE with non-Euclidean distances. Another potential suggestion for future studies would be to explore covariance and HCE models with the Spherical manifold (geodesic distance) for ore deposits covering large parts of the Earth’s surface or lenticular deposits. Acknowledgement The research project is implemented in the framework of H.F.R.I call “Basic research Financing (Horizontal support of all Sciences)” under the National Recovery and Resilience Plan “Greece 2.0” funded by the European Union – NextGenerationEU (H.F.R.I. Project Number: 16537). 16 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 References [1] Curriero, F.C.: On the use of non-euclidean distance measures in geostatistics. Mathematical Geology 38, 907–926 (2006) [2] Theodoridou, P., Varouchakis, E., Karatzas, G.: Spatial analysis of groundwater levels using fuzzy logic and geostatistical tools. Journal of Hydrology 555, 242– 252 (2017) [3] BJK Davis, F.C.: Development and evaluation of geostatistical methods for noneuclidean-based spatial covariance matrices. Mathematical geosciences (2019) [4] Hristopulos, D.T.: Non-separable covariance kernels for spatiotemporal gaussian processes based on a hybrid spectral method and the harmonic oscillator. IEEE Transactions on Information Theory (2023) [5] Agou, V.D., Varouchakis, E.A., Hristopulos, D.T.: Geostatistical analysis of precipitation in the island of Crete (Greece) based on a sparse monitoring network. Environmental Monitoring and Assessment 191(353), 1573–2959 (2019) https://doi.org/10.1007/s10661-019-7462-8 [6] Varouchakis, E.A., Guardiola-Albert, C., Karatzas, G.P.: Spatiotemporal geostatistical analysis of groundwater level in aquifer systems of complex hydrogeology. Water Resources Research 58(3), 2021–029988 (2022) [7] Ver Hoef, J.M.: Kriging models for linear networks and non-euclidean distances: Cautions and solutions. Methods in Ecology and Evolution (2018) [8] Christakos, G.: Random Field Models in Earth Sciences. Academic Press, San Diego (1992) [9] Chil`es, J.P., Delfiner, P.: Geostatistics: Modeling Spatial Uncertainty. Wiley, Hoboken, NJ, USA (2012) [10] Hristopulos, D.T.: Random Fields for Spatial Data Modeling: A Primer for Scientists and Engineers. Springer, Dordrecht, the Netherlands (2020). https: //doi.org/10.1007/978-94-024-1918-4 [11] Pavlides, A., Hristopulos, D.T., Roumpos, C., Agioutantis, Z.: Spatial modeling of lignite energy reserves for exploitation planning and quality control. Energy 93, 1906–1917 (2015) https://doi.org/10.1016/j.energy.2015.10.049 [12] Pavlides, A., Agou, V.D., Hristopulos, D.T.: Non-parametric kernel-based estimation and simulation of precipitation amount. Journal of Hydrology 612, 127988 (2022) [13] Hristopulos, D.T., Pavlides, A., Agou, V.D., Gkafa, P.: Stochastic local interaction model: An alternative to kriging for massive datasets. Mathematical Geosciences 17 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 53(8), 1907–1949 (2021) [14] Agou, V.D., Pavlides, A., Hristopulos, D.T.: Spatial modeling of precipitation based on data-driven warping of gaussian processes. Entropy 24(3), 321 (2022) [15] Pavlides, A., Varouchakis, E., Hristopulos, D.: Geostatistical analysis of groundwater levels in a mining area with three active mines. Hydrogeology Journal 31(6), 1425–1441 (2023) [16] Ghosh, S.: Kernel Smoothing: Principles, Methods and Applications. John Wiley & Sons, Hoboken, NJ, USA (2018) [17] Goovaerts, P.: Geostatistics for Natural Resources Evaluation. Oxford University Press, New York (1997) [18] Emery, X., Arroyo, D., Mery, N.: Twenty-two families of multivariate covariance kernels on spheres, with their spectral representations and sufficient validity conditions. Stochastic Environmental Research and Risk Assessment 36(5), 1447–1467 (2022) [19] Bochner, S., Tenenbaum, M., Pollard, H.: Lectures on Fourier Integrals. Annals of mathematics studies. Princeton University Press, New Jersey (1959) [20] Graczyk, P.: Cram´er theorem on symmetric spaces of noncompact type. Journal of Theoretical Probability 7, 609–613 (1994) [21] Emery, X., Alegr´ıa, A., Arroyo, D.: Covariance models and simulation algorithm for stationary vector random fields on spheres crossed with euclidean spaces. SIAM Journal on Scientific Computing 43(5), 3114–3134 (2021) [22] Armstrong, M.: Common problems seen in variograms. Mathematical Geology 16(3), 305–313 (1984) [23] Elogne, S., Hristopulos, D., Varouchakis, E.: An application of Spartan spatial random fields in environmental mapping: focus on automatic mapping capabilities. Stochastic Environmental Research and Risk Assessment 22(5), 633–646 (2008) https://doi.org/10.1007/s00477-007-0167-5 [24] Hristopulos, D.T.: Spartan Gibbs random field models for geostatistical applications. SIAM Journal of Scientific Computing 24(6), 2125–2162 (2003) [25] Hristopulos, D.T., Elogne, S.: Analytic properties and covariance functions of a new class of generalized Gibbs random fields. IEEE Transactions on Information Theory 53(12), 4667–4679 (2007) [26] Hristopulos, D.T., Elogne, S.N.: Computationally efficient spatial interpolators based on spartan spatial random fields. IEEE Transactions on Signal Processing 57(9), 3475–3487 (2009) https://doi.org/10.1109/tsp.2009.2021450 18 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 [27] Varouchakis, E.A., Hristopulos, D.T.: Improvement of groundwater level prediction in sparsely gauged basins using physical laws and local geographic features as auxiliary variables. Advances in Water Resources 52, 34–49 (2013) https: //doi.org/10.1016/j.advwatres.2012.08.002 [28] Yaglom, A.M.: Correlation Theory of Stationary and Related Random Functions, Volume I: Basic Results vol. 131. Springer, Berlin, Germany (1987) [29] Cressie, N.: Spatial Statistics. John Wiley and Sons, New York (1993) [30] Olea, R.A.: Geostatistics for Engineers and Earth Scientists. Springer, New York, NY (2012) [31] Wackernagel, H.: Multivariate Geostatistics: an Introduction with Applications. Springer, Berlin, Germany (2003) [32] Pavlides, A., Agou, V.D., Hristopulos, D.T.: Non-parametric kernel-based estimation and simulation of precipitation amount. Journal of Hydrology 612, 127988 (2022) https://doi.org/10.1016/j.jhydrol.2022.127988 19 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65