VARIABLE SELECTION IN MULTIVARIATE FUNCTIONAL DATA CLASSIFICATION
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Górecki, Tomasz; Krzyśko, Mirosław; Wołyński, Waldemar Article VARIABLE SELECTION IN MULTIVARIATE FUNCTIONAL DATA CLASSIFICATION Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Górecki, Tomasz; Krzyśko, Mirosław; Wołyński, Waldemar (2019) : VARIABLE SELECTION IN MULTIVARIATE FUNCTIONAL DATA CLASSIFICATION, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, NY, Vol. 20, Iss. 2, pp. 123-138, https://doi.org/10.21307/stattrans-2019-018 This Version is available at: https://hdl.handle.net/10419/207938 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/
STATISTICS IN TRANSITION new series, June 2019 123 STATISTICS IN TRANSITION new series, June 2019 Vol. 20, No. 2, pp. 123–138, DOI 10.21307/stattrans-2019-018 VARIABLE SELECTION IN MULTIVARIATE FUNCTIONAL DATA CLASSIFICATION Tomasz Górecki1,Mirosław Krzy´ sko2, Waldemar Woły´ nski3 ABSTRACT A new variable selection method is considered in the setting of classification with multivariate functional data (Ramsay and Silverman (2005)). The variable selection is a dimensionality reduction method which leads to replace the whole vector process, with a low-dimensional vector still giving a comparable classification error. Various classifiers appropriate for functional data are used. The proposed variable selection method is based on functional distance covariance (dCov) given by Székely and Rizzo (2009, 2012) and the Hilbert-Schmidt Independent Criterion (HSIC) given by Gretton et al. (2005). This method is a modification of the procedure given by Kong et al. (2015). The proposed methodology is illustrated with a real data example. Key words: multivariate functional data, variable selection, dCov, HSIC, classification. 1. Introduction In recent years, much attention has been paid to methods for representing data as functions or curves. Such data are known in the literature as functional data (Ramsay and Silverman (2005), Horváth and Kokoszka (2012)). Applications of functional data can be found in various fields, including medicine, economics, meteorology and many others. In many applications there is a need to use statistical methods for objects characterized by multiple variables observed at many time points (doubly multivariate data). Such data are called multivariate functional data. In this paper we focus on the classification problem for multivariate functional data. In many cases, in the classification procedures, the number of predictors pis significantly greater than the sample size n. Thus, it is natural to assume that only a small number of predictors are relevant to response Y. Various basic classification methods have also been adapted to functional data, such as linear discriminant analysis (Hastie et al. (1995)), logistic regression (Rossi 1Faculty of Mathematics and Computer Science, Adam Mickiewicz University, Poland. E-mail: [email protected]. ORCID ID: https://orcid.org/0000-0002-9969-5257. 2Faculty of Mathematics and Computer Science, Adam Mickiewicz University, Poland. Interfaculty Institute of Mathematics and Statistics, The President Stanisław Wojciechowski State University of Applied Sciences in Kalisz, Poland. E-mail: [email protected]. ORCID ID: https://orcid.org/00000001-0075-4432. 3Faculty of Mathematics and Computer Science, Adam Mickiewicz University, Poland. E-mail: w[email protected]. ORCID ID: https://orcid.org/0000-0002-0777-9163.
124 T. Górecki, M. Krzy´ sko, W. Woły´ nski: Variable selection ... et al. (2002)), penalized optimal scoring (Ando (2009)), knn (Ferraty and Vieu (2003)), SVM (Rossi and Villa (2006)), and neural networks (Rossi et al. (2005)). Moreover, the combining of classifiers has been extended to functional data (Ferraty and Vieu (2009)). Górecki et al. (2016) adapted multivariate regression models to the classification of multivariate functional data. Gretton et al. (2005) defined the measure of dependence between random vectors X X Xand Y Y Ycalled the HilbertSchmidt Independence Criterion (HSIC) and proved that this measure is equal to zero if and only if X X Xand Y Y Yare independent to each other when using universal kernels, such as the Gaussian kernels. Based on the idea of HSIC between two random vectors, we introduced the HSIC between two random processes. Székely et al. (2007), Székely and Rizzo (2009, 2012, 2013) defined the measures of dependence between random vectors: the distance covariance (dCov). These authors showed that for all random variables with finite first moments, dCov generalizes the idea of covariance in two ways. Firstly, this coefficient can be applied when X X Xand Y Y Yare of any dimensions and not only for the simple case where p=q=1. Secondly, dCov is equal to zero if and only if there is independence between the random vectors. Indeed, the distance covariance measures a linear relationship and can be equal to 0 even when the variables are related. Based on the idea of the distance covariance between two random vectors, we introduced the functional distance covariance between two random processes. We select a set of important predictors with a large value of functional distance covariance or functional Hilbert-Schmidt Independent Criterion. Our selection procedure is a modification of the procedure given by Kong et al. (2015). An entirely different approach to the variable selection in functional data classification is presented by Berrendero et al. (2016). It is clear that variable selection has, at least, an advantage when compared with other dimension reduction methods (functional principal component analysis (FPCA), see Górecki et al. (2014), Jacques and Preda (2014), functional partial least squares (FPLS) methodology, see Delaigle and Hall (2012), and other methods) based on general projections: the output of any variable selection method is always directly interpretable in terms of the original variables, provided that the required number dof the selected variables is not too large. The rest of this paper is organized as follows. In Section 2 we present the classification procedures used through the paper. In Section 3 we present the problem of representing functional data by orthonormal basis functions. In Section 4, we define a functional distance covariance. In Section 5 we define a functional HSIC. In Section 6 we propose a variable selection procedure based on the functional distance covariance and on HSIC. In Section 7 we illustrate the proposed methodology through a real data example. We conclude in Section 8. 2. Classifiers The classification problem involves determining a procedure by which a given object can be assigned to one of qpopulations based on observation of pfeatures of that
STATISTICS IN TRANSITION new series, June 2019 125 object. The object being classified can be described by a random pair (X X X,Y), where X X X= (X1,X2,...,Xp)>∈Rpand Y∈ {1,...,q}. An automated classifier can be viewed as a method of estimating the posterior probability of membership in groups. For a given X X X, a reasonable strategy is to assign X X Xto that class with the highest posterior probability. This strategy is called the Bayes’ rule classifier. 2.1. Linear and quadratic discriminant classifiers Now we make the Bayes’ rule classifier more specific by the assumption that all multivariate probability densities are multivariate normal having arbitrary mean vectors and a common covariance matrix. We shall call this model the linear discriminant classifier (LDC). Assuming that class-covariance matrices are different, we obtain quadratic discriminant classifier (QDC). 2.2. Naive Bayes classifier A naive Bayes classifier is a simple probabilistic classifier based on applying Bayes’ theorem with independence assumptions. When dealing with continuous data, a typical assumption is that the continuous values associated with each class are distributed according to a one-dimensional normal distribution or we estimate density by kernel method. 2.3. k-nearest neighbour classifier Most often we do not have sufficient knowledge of the underlying distributions. One of the important nonparametric classifiers is a k-nearest neighbour classifier (kNN classifier). Objects are assigned to the class having the majority in the knearest neighbours in the training set. 2.4. Multinomial logistic regression It is a classification method that generalizes logistic regression to multiclass problem using one vs. all approach. 3. Functional data We now assume that the object being classified is described by a p-dimensional random process X X X= (X1,X2,...,Xp)>∈Lp 2(I), where L2(I)is the Hilbert space of square-integrable functions, and E(X X X) = 0 0 0. Moreover, assume that the kth component of the vector X X Xcan be represented by a finite number of orthonormal basis functions {ϕb} Xk(t) = Bk ∑ b=0 αkbϕb(t),t∈I,k=1,...,p,
126 T. Górecki, M. Krzy´ sko, W. Woły´ nski: Variable selection ... where αk0,αk1,...,αkBkare the unknown coefficients. Let α α α= (α10,...,α1B1,...,αp0,...,αpBp)>∈RK+p,K=B1+···+Bp and Φ Φ Φ(t) = ϕ ϕ ϕ> 1(t)0 0 0... 0 0 0 0 0 0ϕ ϕ ϕ> 2(t)... 0 0 0 ... ... ... ... 0 0 0 0 0 0... ϕ ϕ ϕ> p(t) , where ϕ ϕ ϕk(t) = (ϕ0(t),...,ϕBk(t))>,k=1,..., p. Using the above matrix notation, the process X X Xcan be represented as: X X X(t) = Φ Φ Φ(t)α α α,(1) where E(α α α) = 0 0 0. This means that the realizations of the process X X Xare in finite dimensional subspace of Lp 2(I). We will denote this subspace by Lp 2(I). We can estimate the vectorα α αon the basis of nindependent realizationsx x x1,x x x2,...,x x xn of the random process X X X(functional data). We will denote this estimator by ˆ α α α. Typically data are recorded at discrete moments in time. Let xk j denote an observed value of the feature Xk,k=1,2,...,pat the jth time point tj, where j= 1,2,...,J. Then our data consist of the pJ pairs (tj,xk j). These discrete data can be smoothed by continuous functions xkand Iis a compact set such that tj∈I, for j=1,...,J. Details of the process of transformation of discrete data to functional data can be found in Ramsay and Silverman (2005) or in Górecki et al. (2014). 4. Distance covariance (dCov) For the jointly distributed random process X X X∈Lp 2(I)and the random vector Y Y Y∈Rq, let fX X X,Y Y Y(l l l,m m m) = E{exp[ihl l l,X X Xi+ihm m m,Y Y Yiq]} be the joint characteristic function of (X X X,Y Y Y), where hl l l,X X Xi=ZI l l l0(t)X X X(t)dt and hm m m,Y Y Yi=m m m0Y Y Y. Moreover, we define the marginal characteristic functions of X X XandY Y Yas follows: fX X X(l l l) = fX X X,Y Y Y(l l l,0 0 0)and fY Y Y(m m m) = fX X X,Y Y Y(0 0 0,m m m). Here, for generality, we assume that Y Y Y∈Rq, although the label Yin the classification problem is a random variable, with values in {1,...,q}. Label Y Y Yhas to be transformed into the label vector Y Y Y= (Y1,...,Yq)0, where Yi=1for i=1,...,qif X X X belongs to class i, and 0otherwise.
STATISTICS IN TRANSITION new series, June 2019 127 Now, let us assume that X X X∈Lp 2(I). Then, the process X X Xhas the representation (1). In this case, we may assume (Ramsay and Silverman (2005)) that the vector weight function l l land the process X X Xare in the same space, i.e. the function l l lcan be written in the form l l l(t) = Φ Φ Φ(t)λ λ λ,(2) where λ λ λ∈RK+p. Hence hl l l,X X Xi=ZI l l l0(t)X X X(t)dt =λ λ λ0[ZI Φ Φ Φ0(t)Φ Φ Φ(t)dt]α α α=λ λ λ0α α α, where α α αand λ λ λare vectors occurring in the representations (1) and (2) of the process X X Xand function l l l, and fX X X,Y Y Y(l l l,m m m) = E{exp[iλ λ λ0α α α+im m m0Y Y Y]}=fα α α,Y Y Y(λ λ λ,m m m), where fα α α,Y Y Y(λ λ λ,m m m)is the joint characteristic function of the pair of random vectors (α α α,Y Y Y). On the basis of the idea of distance covariance between two random vectors (Székely et al. (2007)), we can introduce functional distance covariance between random process X X Xand random vector Y Y Y. Definition 1. A nonnegative number dCov(X X X,Y Y Y)defined by dCov(X X X,Y Y Y) = dCov(α α α,Y Y Y), where dCov2(α α α,Y Y Y) = 1 CK+pCqZRK+p+q |fα α α,Y Y Y(λ λ λ,m m m)−fα α α(λ λ λ)fY Y Y(m m m)|2 kλ λ λkK+p+1 K+pkm m mkq+1 q dλ λ λdm m m, and |z|denotes the modulus of z∈C,kλ λ λkK+p,km m mkqthe standard Euclidean norms on the corresponding spaces Vchosen to produce scale free and rotation invariant measure that does not go to zero for dependent random vectors, and Cr=π1 2(r+1) Γ(1 2(r+1)) is half the surface area of the unit sphere in Rr+1, is called a functional distance covariance between the random process X X Xand the random vector Y Y Y. We can estimate functional distance covariance using data set S S S={(ˆ α α α1,y y y1),...,(ˆ α α αn,y y yn)}.
128 T. Górecki, M. Krzy´ sko, W. Woły´ nski: Variable selection ... Let ¯ α α α=1 n n ∑ i=1 ˆ α α αk,¯ y y y=1 n n ∑ i=1 y y yk, ˜ α α αk=ˆ α α αk−¯ α α α,˜ y y yk=y y yk−¯ y y y,k=1,...,n and A A A= (akl ),B B B= (bkl ), ˜ A A A= (Akl ),˜ B B B= (Bkl ), where akl =kˆ α α αk−ˆ α α αlkK+p,bkl =ky y yk−y y ylkq, Akl =k˜ α α αk−˜ α α αlkK+p,Bkl =k˜ y y yk−˜ y y ylkq,k,l=1,...,n. Hence ˜ A A A=H H HA A AH H H,˜ B B B=H H HB B BH H H, where H H H=I I In−1 n1 1 1n1 1 10 n is the centring matrix. On the basis of the result of Székely et al. (2007), we have dCov(S S S) = 1 n2 n ∑ k,l=1 Akl Bkl. 5. Hilbert-Schmidt Independent Criterion (HSIC) Let φ φ φbe a mapping from Lp 2to an inner product feature space H, and ψ ψ ψbe a mapping from Rqto an inner product feature space G. Definition 2. The cross-covariance operator C C CX X X,Y Y Y:G→His a linear operator defined as C C CX X X,Y Y Y=EX X X,Y Y Y[φ φ φ(X X X)⊗ψ ψ ψ(Y Y Y)]−µX X X⊗µY Y Y, for all f∈Hand g∈G, where the tensor product operator f⊗g:G→H,f∈H, g∈G, is defined as (f⊗g)h=fhg,hiG,for all h∈G. This is a generalization of the cross-covariance matrix between random vectors.
STATISTICS IN TRANSITION new series, June 2019 129 Moreover, by the definition of the Hilbert-Schmidt (HS) norm, we can compute the HS norm of f⊗gvia kf⊗gk2 HS =kfk2 Hkgk2 G. Gretton et al. (2005) defined the Hilbert-Schmidt Independence Criterion (HSIC) in the following way: Definition 3. Hilbert-Schmidt Independence Criterion (HSIC) is the squared HilbertSchmidt norm of the cross-covariance operator HSIC(X X X,Y Y Y) = kC C CX X X,Y Y Yk2 HS. Now, let k:Rp×Rp→R be a kernel function on Rp. This raises an interesting question: given a function of two variables k(x x x,x x x0), does there exist a function φ φ φsuch that k(x x x,x x x0) = hφ φ φ(x x x),φ φ φ(x x x0)iH?The answer is provided by Mercer’s theorem (1909), which says, roughly, that if kis positive semidefinite then such a φ φ φexists. Often, we will not know φ φ φ, but a kernel function k, which encodes the inner product in H, instead. Popular positive semi-definite kernel functions on Rpinclude the polynomial kernel of degree d>0,k(x x x,x x x0) = (1+x x x>x x x0)d, the Gaussian kernel k(x x x,x x x0) = exp(−λkx x x− x x x0k2),λ>0, and the Laplace kernel k(x x x,x x x0) = exp(−λkx x x−x x x0k),λ>0. In this paper we use, the Gaussian kernel. A feature mapping φ φ φis centred by subtracting from it its expectation, that is transforming φ φ φ(x x x)to ˜ φ φ φ(x x x) = φ φ φ(x x x)−EX X X[φ φ φ(X X X)]. Centring a positive semi-definite kernel function kconsists in centring in the feature mapping φ φ φassociated to k. Thus, the centred kernel ˜ kassociated to kis defined by ˜ k(x x x,x x x0) = hφ φ φ(x x x)−EX X X[φ φ φ(X X X)],φ φ φ(x x x0)−EX X X0[φ φ φ(X X X0)]i =k(x x x,x x x0)−EX X X[k(X X X,x x x0)]−EX X X0[k(x x x,X X X0)]+EX X X,X X X0[k(X X X,X X X0)], assuming the expectations exist. Here, the expectation is taken over independent copies X X X,X X X0. We see that, ˜ kis also a positive semi-definite kernel. Note also that for a centred kernel ˜ k,EX X X,X X X0[˜ k(X X X,X X X0)] = 0, that is, centring the feature mapping implies centring the kernel function. Let {x x x1,...,x x xn}be a finite subset of Rp. A feature mapping φ φ φis centred by subtracting from it its empirical expectation, i.e. leading to ¯ φ φ φ(x x xi) = φ φ φ(x x xi)−φ φ φ, where φ φ φ=1 n∑n i=1φ φ φ(x x xi). The kernel matrix K K K= (Ki j)associated to the kernel function k and the set {x x x1,...,x x xn}is centred by replacing it with e K K K= ( e Ki j)defined for all i,j=
130 T. Górecki, M. Krzy´ sko, W. Woły´ nski: Variable selection ... 1,2,...,nby ˜ Ki j =Ki j −1 n n ∑ i=1 Ki j −1 n n ∑ j=1 Ki j +1 n2 n ∑ i,j=1 Ki j, where Ki j =k(x x xi,x x xj),i,j=1,...,n. The centred kernel matrix e K K Kis a positive semi-definite matrix. Also, as with the kernel function 1 n2∑n i,j˜ Ki j =0. Let h·,·iFdenote the Frobenius product and k · kFthe Frobenius norm defined for all A A A,B B B∈Rn×nby hA A A,B B BiF=tr(A A A>B B B), kA A AkF= (hA A A,A A AiF)1/2. Then, for any kernel matrix K K K∈Rn×n, the centred kernel matrix e K K Kcan be expressed as follows (Schölkopf et al.(1998)): e K K K=H H HK K KH H H, where H H Hia a centering matrix. Since H H His the idempotent matrix (H H H2=H H H), then we get for any two kernel matrices K K Kand L L Lbased on the subset {x x x1,...,x x xn}of Rpand the subset {y y y1,...,y y yn}of Rq, respectively, he K K K,e L L LiF=hK K K,e L L LiF=he K K K,L L LiF. We may express HSIC in terms of kernel functions (Gretton et al. (2005)): HSIC(X X X,Y Y Y) = EX X X,X0 X0 X0,Y Y Y,Y0 Y0 Y0[k(X X X,X X X0)l(Y Y Y,Y Y Y0)] +EX X X,X0 X0 X0[k(X X X,X X X0)]EY Y Y,Y0 Y0 Y0[l(Y Y Y,Y Y Y0)] −2EX X X,Y Y Y[EX0 X0 X0[k(X X X,X X X0)]EY0 Y0 Y0[l(Y Y Y,Y Y Y0)]]. Here, EX X X,X0 X0 X0,Y Y Y,Y0 Y0 Y0denotes the expectation over independent pairs (X X X,Y Y Y)and (X X X0,Y Y Y0). Let k?:Lp 2(I)×Lp 2(I)→R be a kernel function on Lp 2(I). For the multivariate functional data the Gaussian kernel has the form: k?(x x x,x x x0) = exp(−λkx x x−x x x0k2),λ>0.
STATISTICS IN TRANSITION new series, June 2019 137 KONG, J., WANG, S., WAHBA G., (2015). Using distance covariance for improved variable selection with application to learning genetic risk models, Statistics in Medicine, 34, pp. 1708–1720. KUHN, M., Contributions from Jed Wing, Steve Weston, Andre Williams, Chris Keefer, Allan Engelhardt, Tony Cooper, Zachary Mayer, Brenton Kenkel, the R Core Team, Michael Benesty, Reynald Lescarbeau, Andrew Ziem, Luca Scrucca, Yuan Tang, Can Candan and Tyler Hunt, (2018), caret: Classification and Regression Training. R package version 6.0-80, https://CRAN.Rproject.org/package=caret. R Core Team (2018). R: A language and environment for statistical computing, R Foundation for Statistical Computing, Vienna, Austria, https://www.Rproject.org/. RAMSAY, J. O., SILVERMAN, B.W., (2005). Functional Data Analysis, Springer, New York. RAMSAY, J. O., WICKHAM, H. GRAVES, S., HOOKER, G., (2018). fda: Functional Data Analysis, R package version 2.4.8, https://CRAN.R-project.org/package=fda. RIZZO, M. L., SZÉKELY, G. J., (2018). energy: E-Statistics: Multivariate Inference via the Energy of Data, R package version 1.7-5, https://CRAN.Rproject.org/package=energy. ROSSI, F., DELANNAYC, N., CONAN-GUEZA, B., VERLEYSENC, M., (2005). Representation of functional data in neural networks, Neurocomputing, 64, pp. 183–210. ROSSI, F., VILLA, N., (2006). Support vector machines for functional data classification, Neural Computing, 69, pp. 730–742. ROSSI, N., WANG, X., RAMSAY, J.O., (2002). Nonparametric item response function estimates with EM algorithm, Journal of Educational and Behavioral Statistics, 27, pp. 291–317. SCHÖLKOPF, B., SMOLA, A. J., MÜLLER, K. R., (1998). Nonlinear component analysis as a kernel eigenvalue problem, Neural Computation, 10, pp. 1299– 1319. SZÉKELY, G. J., RIZZO, M. L., BAKIROV, N. K., (2007). Measuring and testing dependence by correlation of distances, The Annals of Statistics, 35 (6), pp. 2769–2794.
138 T. Górecki, M. Krzy´ sko, W. Woły´ nski: Variable selection ... SZÉKELY, G. J., RIZZO, M. L., (2009). Brownian distance covariance, Annals of Applied Statistics, 3 (4), pp. 1236–1265. SZÉKELY, G. J., RIZZO, M. L., (2012). On the uniqueness of distance covariance, Statistical Probability Letters, 82 (12), pp. 2278–2282. SZÉKELY, G. J., RIZZO, M. L., (2013). The distance correlation t-test of independence in high dimension. Journal of Multivariate Analysis, 117, pp. 193–213.