scieee AI-readable full text Open interactive document viewer

Supporting vectors vs. principal components

Márquez, Almudena P.,Mengíbar Rodríguez, Míriam

Abstract

Spanish Government FEDER-UCA18-105867

Full text

http://www.aimspress.com/journal/Math AIMS Mathematics, 8(1): 1937–1958. DOI:10.3934/math.2023100 Received: 10 July 2022 Revised: 29 September 2022 Accepted: 08 October 2022 Published: 26 October 2022 Research article Supporting vectors vs. principal components Almudena P. M´ arquez1,*, Francisco Javier Garc´ ıa-Pacheco1, M´ ıriam Mengibar-Rodr´ ıguez2and Alberto S´ anchez-Alzola3 1Department of Mathematics, College of Engineering, University of Cadiz, 11519 Puerto Real, Spain 2Department of Computer Science and Artificial Intelligence, University of Granada, 18071 Granada, Spain 3Department of Statistics and Operation Research, University of Cadiz, 11519 Puerto Real, Spain *Correspondence: Email: [email protected]. Abstract: Let T:X→Ybe a bounded linear operator between Banach spaces X,Y. A vector x0∈SX in the unit sphere SXof Xis called a supporting vector of Tprovided that kT(x0)k=sup{kT(x)k:kxk= 1}=kTk. Since matrices induce linear operators between finite-dimensional Hilbert spaces, we can consider their supporting vectors. In this manuscript, we unveil the relationship between the principal components of a matrix and its supporting vectors. Applications of our results to real-life problems are provided. Keywords: bounded linear operator; Hilbert space; mean operator; principal components; supporting vector Mathematics Subject Classification: 51F30, 54E35, 54E45 1. Introduction Supporting Vector Analysis (SVA) is a relatively recent technique that allows one to solve analytically many real-life problems that used to be tackled by means of Heuristic methods. The lack of mathematical formalism of Heuristic methods resulted many times in unpredictable solutions, that is, mathematical solutions whose real-life interpretations make no sense. Supporting vectors came into play to overcome this issue. This way, supporting vectors were used in a successful way to solve multiobjective optimization problems coming from different disciplines, such as Bioengineering, Physics, and Statistics [4, 6–9, 15, 22], improving considerably the results achieved by other methods like, for instance, Heuristic techniques [10,11,20, 21]. In [4, 6, 15], it was proven that Singular Value Decomposition (SVD) can be seen as a particular case of SVA. This fact triggered the new trend of restating Statistical notions from the perspective of 1938 Functional Analysis and Operator Theory. The main objective of this manuscript is to study Principal Component Analysis (PCA) by means of SVA. 2. Materials and methods We will review several basic notions from Operator Theory that will turn out to be crucial for the development of this manuscript. 2.1. Centering and standardizing If x=(x1,...,xn)∈Rn, then the mean of xis defined as x:=1 nPn i=1xi, and its standard deviation is given by sx:=q1 nPn i=1(xi−x)2. Notice that √nsx=x−x2,(2.1.1) where x:=(x,n . . ., x)denotes the constant vector of term x(in general, if a∈R, then a:=(a,n . . ., a) denotes the constant vector of term a). We say that x∈Rnis centered provided that x=0, and it is standardized provided that x=0 and sx=1. In the latter situation, kxk2=√n, in view of (2.1.1). The subset of centered vectors of Rnis usually denoted by cen(Rn), that is, cen(Rn) :={x∈Rn:x=0}.The subset of standardized vectors of Rnis usually denoted by stan(Rn), that is, stan(Rn) :={x∈Rn:x=0 and sx=1}. According to (2.1.1), stan (Rn)⊆√nS`n 2, where S`n 2stands for the unit sphere of `n 2:=(Rn,k·k2). In Topology, S`n 2is denoted as Sn−1. 2.2. Principal component analysis The covariance of two vectors x,y∈Rnis defined as sx,y:=1 n n X i=1 (xi−x) (yi−y). Notice that sx,x=s2 x, that is, the variance of x. The covariance matrix of a given matrix A∈Mm×nis defined by sa1,...,an:=sai,aji,j=1,...,n, where a1,...,anstand for the column vectors of A. Consider a matrix A∈Mm×n. The principal components of Aare defined as Ax1,...,Axn, where {x1,...,xn}is an ordered orthonormal basis of eigenvectors of sa1,...,an, sorting the eigenvalues of sa1,...,an decreasingly. We refer the reader to [24] for a wider perspective on PCA. Interesting applications of PCA to certain Engineering fields, such as video processing and Big Data, have been provided in [3,12]. 2.3. Supporting vector analysis Let X,Ybe Banach spaces. Let T:X→Ybe a bounded linear operator. The operator norm of Tis given by kTk:=sup{kT(x)k:kxk=1}.(2.3.1) AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1939 The vector space CL(X,Y) of continuous linear operators from Xto Ybecomes a Banach space when endowed with the operator norm (2.3.1). In the case X=Y,CL(X,Y) is simply denoted as CL(X). If Y=K(Ror C), then CL(X,Y) is denoted as X∗, that is, the dual space of X. It is also common to denote CL(X,Y) by B(X,Y) and CL(X) by B(X). The supporting vector notion was formally posed for the first time in [5]. However, this concept can be found implicitly and scattered throughout the literature of Banach Space Theory [1,2, 18,19]. The set of supporting vectors of a bounded linear operator T:X→Ybetween Banach spaces X,Y is defined by suppv(T) :={x∈SX:kT(x)k=kTk} =arg max kxk=1kT(x)k.(2.3.2) Here, SXstands for the unit sphere of X, and BXdenotes the (closed) unit ball of X. In the infinitedimensional setting, it may occur that (2.3.2) is empty. Note that suppv(T)=suppv(λT) for all λ∈ K\ {0}, and suppv(T)=SKsuppv(T), where K=Ror C. For a topological and geometrical analysis of the above set, we strongly refer the reader to [13,14,23]. For linear functionals, a special subset of supporting vectors is worth regarding. Consider a continuous linear functional f∈X∗in the dual X∗of a Banach space X. We define the set of 1supporting vectors of fby suppv1(f) :={x∈SX:f(x)=kfk}.(2.3.3) Notice that 1-supporting vectors are particular cases of supporting vectors; in other words, suppv1(f)⊆ suppv( f). In the upcoming sections, 1-supporting vectors will be very much relied on. The following remark highlights a standard geometrical property satisfied by 1-supporting vectors. Remark 2.1. Consider a Banach space X and a nonzero linear functional f ∈X∗\ {0}. For every x,y∈suppv1(f)and every λ∈[0,1], we have that λx+(1 −λ)y∈suppv1(f), that is, suppv1(f)is a convex subset of the unit sphere SXof X. A direct consequence of Remark 2.1 is that suppv1(f) is either empty or a singleton in strictly convex Banach spaces, like, for instance, Hilbert spaces. 2.4. Hilbert space theory Representation Theory is one of the most important theories in Mathematics. A major result in Representation Theory is undoubtedly the Riesz Representation Theorem. This is a key result in Functional Analysis and is crucial for working with self-adjoint operators on Hilbert spaces. Riesz Representation Theorem. In a Hilbert space H, for every h∗∈H∗there exists a unique h ∈H satisfying h∗=(·|h). This assignment between H and H∗is a surjective linear isometry. In view of Remark 2.1 and under the settings of the Riesz Representation Theorem, for every h∈ H\{0}, we have that suppv1(h∗)=nh khko, that is, h khkis the only 1-supporting vector of h∗. On the other hand, the orthogonal subspace of a closed subspace Vof a Hilbert space His denoted by V⊥. The orthogonal projection of Honto Vis usually denoted by pV. Observe that H=V⊕2V⊥, that is, for each h∈H, h=pV(h)+pV⊥(h),and khk2=kpV(h)k2+kpV⊥(h)k2.(2.4.1) AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1940 The adjoint of a bounded linear operator T:H→Kbetween Hilbert spaces H,Kis defined as the unique bounded linear operator T∗:K→Hsatisfying (T(h)|k)=(h|T∗(k))for each h∈Hand each k∈K. Basic properties satisfied by the adjoint operator are kT∗k=kTk,(T∗)∗=T, (T+S)∗=T∗+S∗, (T◦S)∗=S∗◦T∗and (λT)∗=λT∗. Lemma 2.2. Every bounded linear operator T :H→K between Hilbert spaces H,K verifies that cl(T(H)) =ker (T∗)⊥and T(H)⊥=cl(T(H))⊥=ker (T∗). Proof. First off, it is a trivial observation that T(H)⊥=cl(T(H))⊥. Fix an arbitrary h∈H. For every k∈ker (T∗), (T(h)|k)=(h|T∗(k))=(h|0) =0. This shows that T(h)∈ker (T∗)⊥. The arbitrariness of h∈Hmeans that T(H)⊆ker (T∗)⊥. By taking orthogonal complements, we obtain that ker (T∗)⊆T(H)⊥. Conversely, fix an arbitrary k∈T(H)⊥. For every h∈H, 0 =(T(h)|k)=(h|T∗(k)). This shows that T∗(k)=0, and hence k∈ker (T∗). The arbitrariness of k∈T(H)⊥means that T(H)⊥⊆ker (T∗). By taking orthogonal complements, we finally obtain that ker (T∗)⊥⊆cl(T(H)).  A bounded operator T∈B(H) is said to be self-adjoint provided that T∗=T. If His complex, then T∈B(H) is self-adjoint if and only if (T(h)|h)∈Rfor all h∈H. A self-adjoint operator is called positive provided that (T(h)|h)≥0 for all h∈H. For every T∈B(H), σ(T) :={λ∈C:λI−T<U(B(H))}is the spectrum of T, where U(B(H))is the multiplicative group of invertible operators on H. Among other spectral properties, the spectrum is compact and nonempty, and kTk ≥ max |σ(T)|. A special subset of the spectrum, called the point spectrum, σp(T) :={λ∈C: ker(λI−T),{0}},whose elements are called the eigenvalues of T, will be very much employed. Note that σp(T)⊆σ(T). Furthermore, for each λ∈σp(T), VT(λ) := {h∈H:T(h)=λh}stands for the subspace of eigenvectors associated with λ. In case there is no confusion with T, we will simply denote VT(λ) by V(λ). Suppose next that kTkis an eigenvalue of T, that is, kTk ∈ σp(T). In this situation, since kTk ≥ max |σ(T)|, we conclude that kTkis the maximum of |σ(T)|; in other words, kTk=max |σ(T)|. In this case, we write kTk=λmax(T). Observe also that V(kTk)∩SX⊆suppv(T). Indeed, if x∈V(kTk)∩SX, then T(x)=kTkx, so kT(x)k=kTk, and hence x∈suppv(T). Nevertheless, in general, kTk<σp(T), unless, for instance, Tis compact, self-adjoint and positive. This is why we have to rely on the adjoint T∗and on the strongly positive operator T∗◦T. It is straightforward to check that the eigenvalues of a self-adjoint operator are real, and the eigenvalues of a self-adjoint positive operator are positive. When Tis compact, it holds that T∗◦Tis compact, self-adjoint and positive. The following result, on which we will strongly rely later on, can be found in [6, Theorem 4], which is itself a refinement of [14, Theorem 9]. Theorem 2.3. Let H,K be Hilbert spaces. Let T ∈B(H,K). Then, kTk2=kT∗◦Tk, and suppv(T)⊆ suppv (T∗◦T). Furthermore, suppv(T),∅if and only if kT∗◦Tk ∈ σp(T∗◦T). In this situation, kTk=√λmax (T∗◦T), and suppv(T)=V(λmax(T∗◦T))∩SH. 3. Results In this section, we will state and prove all the novel theorems of this work. This section is divided into four subsections, in which we will deal with supporting vectors, principal components and the AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1941 topological structure of the subsets of centered and standardized vectors. 3.1. Topological structure of the subsets of centered and standardized vectors This subsection begins unveiling the topological structure of the subsets cen(Rn) and stan(Rn) of centered and standardized vectors, respectively. Notice that cen(R)={0}and stan(R)=∅. Theorem 3.1. If n ≥1, then cen(Rn)is linearly isometric, and hence homeomorphic, to `n−1 2. If n ≥2, then stan(Rn)is linearly isometric, and hence homeomorphic, to Sn−2. Proof. Notice that x+y=x+y, and tx =tx for all x∈Rnand all t∈R. As a consequence, ·:Rn→R x7→ x=1 nPn i=1xi (3.1.1) is a linear functional (usually called the mean functional). Next, simply observe that cen(Rn)=ker (·). Thus, by bearing in mind (2.1.1), we immediately obtain that stan(Rn)=ker (·)∩√nS`n 2=√nSker(·). In other words, stan(Rn) is precisely a multiple of the unit sphere of the Hilbert subspace ker (·)of `n 2. Since dim (ker (·)) =n−1, we have that ker (·)is linearly isometric to `n−1 2. As a consequence, the unit sphere of ker (·),1 √nstan(Rn), is linearly isometric to the unit sphere of `n−1 2,S`n−1 2=Sn−2. The following lemma aims at computing the norm of the mean functional as an element of the dual space of `n 2as well as its only 1-supporting vector. Lemma 3.2. In `n 2∗,k·k=1 √nand suppv1(·)=n1 √no. Proof. H¨ older’s Inequality ensures that |x|= 1 n n X i=1 xi≤ n X i=1 1 n|xi| ≤  n X i=1 1 n2 1 2 n X i=1 x2 i 1 2 =1 √nkxk2 for every x∈`n 2. This shows that k·k≤1 √n. On the other hand, k1k2=√n, that is, 1 √n∈S`n 2. Finally, 1 √n=1 n n X i=1 1 √n=1 √n. As a consequence, k·k=1 √nand suppv1(·)=n1 √no. As a direct consequence of Lemma 3.2, the standard deviation of a vector x∈Rncan be rewritten as the following (2.1.1): sx=1 √nx−x2=k·kx−x2.(3.1.2) By bearing in mind the Riesz Representation Theorem, Lemma 3.2 assures that if h:=1 n, then h∗=·in H:=`n 2. Notice that h:=1 n=1 n,n . . ., 1 ncan be seen as a (finite) convex series. This fact will help us generalize these concepts to an infinite-dimensional separable Hilbert space setting later on in the Discussion. AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1942 Definition 3.3. For n ≥1, the mean operator is defined by •:Rn→Rn x7→ x=(x,n . . ., x),(3.1.3) and the centering operator is defined as cen : Rn→Rn x7→ cen(x) :=x−x=(x1−x,n . . ., xn−x).(3.1.4) For n ≥2, the standardizing operator is defined as stan : Rn\Rn1→Rn x7→ stan(x) :=√nx−x kx−xk2.(3.1.5) It is a trivial observation that stan(x)=√ncen(x) kcen(x)k2 =cen(x) k·kkcen(x)k2 (3.1.6) for all x∈Rn. Recall that a projection on a Banach space Xis a continuous, linear, and idempotent map P:X→X. Its dual operator P∗:X∗→X∗is also a projection. The complementary projection of Pis defined as IX−P, which is also a projection. Every non-zero projection has norm greater than or equal to 1. A 1-projection is a projection of norm 1 (also called a contractive projection), and a (1,1)-projection is a 1-projection whose complementary projection is also a 1-projection (also called bicontractive). Orthogonal projections in Hilbert spaces are the most representative examples of bicontractive projections. The final result of this first subsection serves to show that both the mean operator and the centering operator are complementary projections to each other of norm 1, that is, bicontractive, for the Euclidean norm. Even more, the mean operator and the centering operator are orthogonal projections to each other. This will allow us to directly obtain the K¨ onig-Huygens Theorem (3.1.7) as a direct consequence of the Pythagorean Theorem in Hilbert spaces. The K¨ onig-Huygens Theorem provides the classical decomposition of the 2-norm of a vector of Rnin terms of the mean and the standard deviation. Theorem 3.4. The centering operator is a linear projection on Rnwhose kernel is ker(cen) =Rn1, whose range is cen(Rn)=ker (·), and whose complementary projection is the mean operator. Furthermore, if we consider Rnendowed with the Euclidean norm, then k•k =kcenk=1, and • and cen are complementary orthogonal projections. As a consequence, for every x ∈Rn, kxk2 2=kxk2 2+kx−xk2 2=n|x|2+s2 x.(3.1.7) Proof. In the first place, notice that •is clearly linear, since •=·1; in other words, x=(x,n . . ., x)= x(1,n . . ., 1)=x1for all x∈Rn. As a consequence, cen is linear as well because cen =IRn−•, that is, the difference of two linear operators on Rn. On the other hand, observe that if x∈Rnis a constant vector, that is, x=t1=tfor some t∈R, then cen(x)=0. Conversely, if cen(x)=0, then x=x=x1, and hence x∈Rn1is a constant vector. This shows that ker(cen) =Rn1. From Theorem 3.1, we already AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1943 know that cen(Rn)=ker (·). Next, let us prove that cen is a projection. Fix an arbitrary x∈Rn. Notice that cen (cen(x))=cen x−x=cen (x)−cen x=x−x−0=x−x=cen (x). Since •=IRn−cen, we conclude that the mean operator is the complementary projection to the centering operator. Finally, let us compute k•k and kcenk. In accordance with Lemma 3.2, for every x∈Rnwe have that kxk2=√n|x| ≤ √n1 √nkxk2=kxk2, meaning that k•k ≤ 1. Next,  1 √n2 = 1 √n,n . . ., 1 √n!2 =√n1 √n=1. This shows that k•k =1. In order to prove that •and cen are orthogonal projections, it only suffices to realize that R1and cen(Rn) are orthogonal subspaces. Indeed, for every t∈Rand every x∈Rnwith x=0, we have that (t1|x)=t(1|x)=tPn i=1xi=tnx =0, and hence R1⊆cen(Rn)⊥, or equivalently, cen(Rn)⊆(R1)⊥. Furthermore, the Pythagorean Theorem in `n 2allows that kt1+xk2 2=kt1k2 2+kxk2 2 for every t∈Rand every x∈Rnwith x=0. Next, if y∈(R1)⊥, then 0 =1 n(1|y)=1 nPn i=1yi=y, meaning that y∈ker (·)=cen(Rn). As a consequence, cen(Rn)⊥=R1, resulting in kxk2 2=kxk2 2+kx−xk2 2=n|x|2+s2 x for every x∈Rn. It only remains to show that kcenk=1, but this is a direct consequence of the fact that cen is an orthogonal projection.  3.2. Supporting vectors and first principal component In this subsection, we will provide sufficient conditions for the supporting vectors to coincide with the first principal component. First, we will need some definitions. Definition 3.5. A matrix is said to be centered (standardized) provided that all of its column vectors are centered (standardized) vectors. The following lemma displays a simple characterization of centered matrices. Lemma 3.6. Let A ∈Mm×n(R). The following conditions are equivalent: 1) A is centered. 2) Ax is centered for all x ∈Rn. Proof. Let {e1,...,en}denote the canonical basis of Rn. Notice that the columns of Aare precisely Aej for j=1,...,n. Suppose first that Ais centered. Fix an arbitrary x∈Rn. The linearity of the mean functional allows that Ax = n X j=1 xjAej= n X j=1 xjAej=0. AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1944 As a consequence, Ax is centered for all x∈Rn. Conversely, suppose now that Ax is centered for all x∈Rn. In particular, Aej=0 for j=1,...,n, meaning that the columns of Aare centered vectors, that is, Ais centered.  The following characterization of centered matrices is a bit more sophisticated. First, a technical lemma is needed. Lemma 3.7. Consider x,y∈Rm. Then, 1) msx,y=x•y if and only if either x or y is centered. 2) x is standardized if and only if x is centered and x •x=m. Proof. 1) Let us observe that msx,y= m X i=1 (xi−x) (yi−y) = m X i=1 xiyi− m X i=1 xyi− m X i=1 xiy+ m X i=1 x y = m X i=1 xiyi−x m X i=1 yi−y m X i=1 xi+x y m X i=1 1 = m X i=1 xiyi−xmy −ymx +x y =x•y−mx y. As a consequence, msx,y=x•yif and only if either xor yis centered. 2) By definition, xis standardized if and only if xis centered and sx=1. We know that then kx−xk2=√msx. In the context of centered vectors, the previous expression becomes √x•x= kxk2=√msx. Therefore, xis standardized if and only if xis centered and x•x=m.  Recall that the covariance matrix of a given matrix A∈Mm×nis defined by sa1,...,an:=sai,aji,j=1,...,n, where a1,...,anstand for the column vectors of A. Also, recall that if Bis a square matrix, then diag(B) stands for the diagonal of B. Proposition 3.8. Let A ∈Mm×n(R). The following conditions are equivalent: 1) A is centered. 2) AtA=msa1,...,an. 3) diag AtA=diag msa1,...,an. AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1945 Proof. Notice that AtA=ai•aji,j=1,...,n. Suppose first that Ais centered. In view of Lemma 3.7, msai,aj=ai•ajfor all i,j∈ {1,...,n}, meaning that AtA=msa1,...,an. Next, if AtA=msa1,...,an, then we trivially have that diag AtA=diag msa1,...,an. Finally, assume that diag AtA=diag msa1,...,an. Then, msai,ai=ai•aifor all i=1,...,n. In accordance with Lemma 3.7, aiis centered for all i=1,...,n, meaning that Ais centered.  As an immediate consequence of Lemma 3.7 and Proposition 3.8, we obtain the following corollary. Corollary 3.9. Let A ∈Mm×n(R). The following conditions are equivalent: 1) A is standardized. 2) AtA=msa1,...,anand diag AtA=(m,n . . ., m). Finally, we have gathered all the necessary tools to prove our main results of this subsection. Simply keep in mind the small observation that if B∈Mn(R), then σp(αB)=ασp(B) and VαB(αλ)=VB(λ) for all α∈R\{0}and all λ∈σp(B). Theorem 3.10. Let A ∈Mm×n(R). If A is centered, then suppv(A)=nx∈S`n 2:Ax is the first principal component of Ao, where A is seen as a linear operator A:`n 2→`m 2 x7→ Ax.(3.2.1) Proof. According to Theorem 2.3, suppv(A)=V(λmax(A∗◦A))∩S`n 2. Notice that the adjoint of A, A∗, coincides with its transpose, At. On the other hand, since Ais centered, Proposition 3.8 allows that AtA=msa1,...,an. Finally, suppv(A)=VA∗◦A(λmax(A∗◦A))∩S`n 2 =VAtAλmax(AtA)∩S`n 2 =Vmsa1,...,anλmax(msa1,...,an)∩S`n 2 =Vmsa1,...,anmλmax(sa1,...,an)∩S`n 2 =Vsa1,...,anλmax(sa1,...,an)∩S`n 2 =nx∈S`n 2:Ax is the first principal component of Ao.  We will conclude this subsection with an example of a centered matrix Awhose last principal component has at least two dimensions. Remark 3.11. Let T :H→K be a bounded linear operator between Hilbert spaces H,K. Then, ker(T)⊆ker (T∗◦T). As a consequence, if 0∈σp(T), then 0∈σp(T∗◦T). Notice that if T∈B(H) is a self-adjoint positive operator on a Hilbert space Hsuch that ker(T), {0}, then 0 is the minimum of σp(T) since all the eigenvalues of Tare positive. AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1952 4. Discussion We will discuss how to transport the mean functional, the mean operator and the centering operator to abstract settings, such as Hilbert spaces, Banach algebras and probability spaces. 4.1. Generalization to Hilbert spaces As mentioned in Lemma 3.2, in view of the Riesz Representation Theorem, if h0:=1 n=1 n,n . . ., 1 n, then h∗ 0=·is precisely the mean functional in H:=`n 2. Actually, kh0k2=1 √n=kh∗ 0k. At this stage, the key is to realize that h0can be seen as a (finite) convex series. Recall that a convex series is a convergent series P∞ n=1tnsuch that P∞ n=1tn=1 and tn≥0 for all n∈N. Notice that if P∞ n=1tnis a convex series, then P∞ n=1t2 n≤P∞ n=1tn=1, and hence (tn)n∈N∈`2. Now, we are in the right position to define the notions of mean and standard deviation on separable Hilbert spaces. Let Hbe a separable Hilbert space with an orthonormal basis (en)n∈N. Fix a convex series P∞ n=1tn. Let h∈Hand write h=P∞ n=1(h|en)en. The mean of h, with respect to (tn)n∈N, is defined as h:=∞ X n=1 tn(h|en).(4.1.1) The mean functional, with respect to (tn)n∈N, is given by ·:H→K h7→ h:=P∞ n=1tn(h|en).(4.1.2) By relying on H¨ older’s Inequality, it is not hard to check that the mean functional is an element of H∗ whose norm is precisely pP∞ n=1t2 n=k(tn)n∈Nk2. In accordance with the Riesz Representation Theorem, there exists h0∈Hsuch that (h|h0)=hfor all h∈H. If we let h=P∞ n=1(h|en)en, then we obtain that ∞ X n=1 (h|en) (en|h0)= ∞ X n=1 (h|en)en h0 =∞ X n=1 tn(h|en).(4.1.3) By taking h=enfor every n∈Nin (4.1.3), we conclude that (en|h0)=tnfor every n∈N. In particular, h0=P∞ n=1tnen. As expected, kh0k=k(tn)n∈Nk2=v t∞ X n=1 t2 n=k·k. In view of the Riesz Representation Theorem, h∗ 0=(·|h0)=·. Notice that, in order to conclude that h∗ 0=(·|h0)=·, the Riesz Representation Theorem is really not needed. Let us get back for a second to `n 2. For every x∈`n 2, x=(x,n . . ., x)=x1=x 1 √n 1 n 1 √n =· 1 √n (x) 1 n 1 √n . Then, going back to a general separable Hilbert space H, the mean operator is defined as H→H h7→ h:=h∗ 0 kh∗ 0k(h)h0 kh0k.(4.1.4) AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1953 It is not hard to check that the mean operator is an orthogonal projection on Hwhose complementary projection is, precisely, the centering operator: H→H h7→ cen(h) :=h−h.(4.1.5) 4.2. Generalization to Banach algebras A Banach algebra is a real or complex algebra Aendowed with a complete vector norm that is also a ring norm, that is, kabk≤kakkbkfor all a,b∈A. We say that Ais unital if it is unitary, that is, it has a unity 1∈A, and k1k=1. In this situation, according to the Hahn-Banach Theorem, there exists a continuous linear functional, which we will denote by 1∗∈A∗, such that k1∗k=1 and 1∗(1)=1. Then, the mean functional is precisely 1∗, that is, defined by A→K a7→ a:=1∗(a).(4.2.1) The mean operator is defined as A→A a7→ a:=1∗(a)1.(4.2.2) Finally, the centering operator is defined as A→A a7→ cen(a) :=a−a.(4.2.3) Following (3.1.2), the variance of an element a∈Acan be defined by s2 a:=1∗a−a2=1∗(a2)−1∗(a)2.(4.2.4) Notice that this generalization to unital Banach algebras presents a weakness: The existence of the functional 1∗∈A∗is guaranteed by the Hahn-Banach Theorem, but its uniqueness is not guaranteed. In fact, in many unital Banach algebras, such as `∞(Λ) for instance, 1is not a smooth point of B`∞(Λ)[16, Theorem 2.9], and thus there are infinitely many functionals of norm 1 attaining their norm at 1. It seems not trivial to overcome this issue. Maybe an option is to try to renorm equivalently the unital Banach algebra in such a way that 1becomes a smooth point of the new unit ball of the algebra or, at least, to find another smooth point in the unit ball. According to [16, Theorem 2.1], the canonical unit vector eλis a smooth point of B`∞(Λ)for each λ∈Λ. Another possibility may rely on constructing the mean functional in a C∗-algebra. Recall that a C∗-algebra is a Banach algebra Aendowed with an antimultiplicative and antilinear involution ∗:A→Asatisfying that kaa∗k=kak2for all a∈A. 4.3. Generalization to probability spaces A probability space is a 3-tuple (Ω,Σ,P) where (Ω,Σ) is a measurable space and P:Σ→[0,1] is a probability measure, that is, a countably additive positive measure such that P(Ω)=1. If Xis a Banach space, by L1((Ω,Σ,P),X) we denote the Banach space of all absolutely integrable functions, that is, L1((Ω,Σ,P),X) :=(f∈XΩ:fis measurable and ZΩkfkdP<∞), AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1954 endowed with the norm kfk1:=ZΩkfkdP. For each f∈L1((Ω,Σ,P),X), the mean of fis defined as µ(f) :=ZΩ f dP.(4.3.1) The mean functional is given by µ:L1((Ω,Σ,P),X)→X f7→ µ(f) :=RΩf dP.(4.3.2) Notice that the mean functional is actually an operator, but we keep calling it functional not to mistake it with the mean operator, which will be defined next. It can be easily shown that the mean functional has norm equal to 1. Indeed, kµ(f)k=ZΩ f dP≤ZΩkfkdP=kfk1 for every f∈L1((Ω,Σ,P),X), meaning that kµk ≤ 1. Now, if we choose any x∈SX, then xχΩ∈ SL1((Ω,Σ,P),X), and µ(xχΩ)=ZΩ xχΩdP=xP(Ω)=x. Hence, kµ(xχΩ)k=ZΩ xχΩdP =kxkP(Ω)=1=kxχΩk1. This proves that kµk=1 and xχΩ∈suppv(µ). The mean operator is then defined as L1((Ω,Σ,P),X)→L1((Ω,Σ,P),X) f7→ µ(f)χΩ,(4.3.3) and the centering operator is L1((Ω,Σ,P),X)→L1((Ω,Σ,P),X) f7→ f−µ(f)χΩ.(4.3.4) Finally, if Xis a unital Banach algebra, then the natural way of defining the variance of f∈ L1((Ω,Σ,P),X) is σ(f) :=µ(f−µ(f)χΩ)2=ZΩ (f−µ(f)χΩ)2dP.(4.3.5) 5. Conclusions It is well known in the literature of the Geometry of Banach Spaces that Hilbert spaces are transitive Banach spaces, meaning that every two elements of the unit sphere of a Hilbert space can be transported one into another by means of a surjective linear isometry. This fact confers Hilbert spaces with a AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1955 certain freedom when it comes to choosing the convex series that defines the mean functional, the mean operator, the centering operator and the standard deviation. Theorem 3.1, Lemma 3.2 and Theorem 3.4 can be transported to the more general scope provided by separable Hilbert spaces discussed in the previous section. In fact, in the Discussion we unveiled how to extend the mean functional, the mean operator, the centering operator and the variance to spaces of absolutely integrable functions defined on a probability space and valued on a unital Banach algebra. On the other hand, the study of Principal Component Analysis through Supporting Vector Analysis is a revolutionary trend that allows one to look at these Statistical concepts from a Functional Analysis viewpoint, which is more general and works in infinite dimensional environments, making possible applications in very specific settings such as, for instance, Quantum Mechanical Systems. Acknowledgments This work has been partially supported by the Research Grant PGC-101514-B-I00 awarded by the Ministry of Science, Innovation and Universities of Spain and co-financed by the 2014-2020 ERDF Operational Programme and by the Department of Economy, Knowledge, Business and University of the Regional Government of Andalusia under Project reference FEDER-UCA18-105867. The APCs have been paid by the Department of Mathematics of the University of Cadiz. Conflict of interest The authors declare that they have no conflict of interest. References 1. E. Bishop, R. R. Phelps, A proof that every Banach space is subreflexive, Bull. Amer. Math. Soc., 67 (1961), 97–98. https://doi.org/10.1090/S0002-9904-1961-10514-4 2. E. Bishop, R. R. Phelps, The support functionals of a convex set, In: Proceedings of Symposia in Pure Mathematics, Vol. VII, Providence, R.I.: Amer. Math. Soc., 1963, 27–35. 3. T. Bouwmans, S. Javed, H. Zhang, Z. Lin, R. Otazo, On the applications of robust PCA in image and video processing, Proc. IEEE,106 (2018), 1427–1457. https://doi.org/10.1109/JPROC.2018.2853589 4. C. Cobos-S´ anchez, F. J. Garcia-Pacheco, J. M. Guerrero Rodriguez, J. R. Hill, An inverse boundary element method computational framework for designing optimal TMS coils, Eng. Anal. Bound. Elem.,88 (2018), 156–169. https://doi.org/10.1016/j.enganabound.2017.11.002 5. C. Cobos-S´ anchez, F. J. Garc´ ıa-Pacheco, S. Moreno-Pulido, S. S´ aez-Mart´ ınez, Supporting vectors of continuous linear operators, Ann. Funct. Anal.,8(2017), 520–530. https://doi.org/10.1215/20088752-2017-0016 6. C. Cobos-S´ anchez, J. A. Vilchez-Membrilla, A. Campos-Jim´ enez, F. J. Garc´ ıa-Pacheco, Pareto optimality for multioptimization of continuous linear operators, Symmetry,13 (2021), 661. https://doi.org/10.3390/sym13040661 AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1956 7. C. Cobos-S´ anchez, M. R. Cabello, ´ A. Q. Oloz´ abal, M. F. Pantoja, Design of TMS coils with reduced lorentz forces: application to concurrent TMS-fMRI, J. Neural Eng.,17 (2020), 016056. https://doi.org/10.1088/1741-2552/ab4ba2 8. C. Cobos-S´ anchez, J. J. J. Garc´ ıa, M. R. Cabello, M. F. Pantoja, Design of coils for lateralized TMS on mice, J. Neural Eng.,17 (2020), 036007. https://doi.org/10.1088/1741-2552/ab89fe 9. C. Cobos-S´ anchez, F. J. Garcia-Pacheco, J. M. Guerrero-Rodriguez, L. Garcia-Barrachina, Solving an IBEM with supporting vector analysis to design quiet TMS coils, Eng. Anal. Bound. Elem.,117 (2020), 1–12. https://doi.org/10.1016/j.enganabound.2020.04.013 10. C. Cobos-S´ anchez, J. M. Guerrero-Rodriguez, ´ A. Q. Oloz´ abal, D. Blanco-Navarro, Novel TMS coils designed using an inverse boundary element method, Phys. Med. Biol.,62 (2016), 73–90. https://doi.org/10.1088/1361-6560/62/1/73 11. C. M. Epstein, E. Wassermann, U. Ziemann, Oxford Handbook of Transcranial Stimulation, New York: Oxford University Press, 2008. https://doi.org/10.1093/oxfordhb/9780198568926.001.0001 12. J. Fan, Q. Sun, W.-X. Zhou, Z. Zhu, Principal component analysis for big data, Wiley StatsRef: Statistics Reference Online, in press. https://doi.org/10.1002/9781118445112.stat08122 13. F. J. Garc´ ıa-Pacheco, E. Naranjo-Guerra, Supporting vectors of continuous linear projections, International Journal of Functional Analysis, Operator Theory and Applications,9(2017), 85– 95. 14. F. J. Garc´ ıa-Pacheco, Lineability of the set of supporting vectors, RACSAM,115 (2021), 41, https://doi.org/10.1007/s13398-020-00981-6 15. F. J. Garcia-Pacheco, C. Cobos-S´ anchez, S. Moreno-Pulido, A. Sanchez-Alzola, Exact solutions to maxkxk=1P∞ i=1kTi(x)k2with applications to Physics, Bioengineering and Statistics, Commun. Nonlinear Sci. Numer. Simul.,82 (2020), 105054. https://doi.org/10.1016/j.cnsns.2019.105054 16. F.-K. Garsiya-Pacheko, The cardinality of the set Λdetermines the geometry of the spaces B`∞(Λ)and B`∞(Λ)∗, (Russian), Funktsional. Anal. i Prilozhen.,52 (2018), 62–71. https://doi.org/10.4213/faa3534 17. Instituto de Estad´ ıstica y Cartograf´ ıa de Andaluc´ ıa. Available from: https://www. juntadeandalucia.es/institutodeestadisticaycartografia. 18. R. C. James, Characterizations of reflexivity, Stud. Math.,23 (1964), 205–216. https://doi.org/10.4064/sm-23-3-205-216 19. J. Lindenstrauss, On operators which attain their norm, Israel J. Math.,1(1963), 139–148. https://doi.org/10.1007/BF02759700 20. L. Marin, H. Power, R. W. Bowtell, C. Cobos-S´ anchez, A. A. Becker, P. Glover, et al., Numerical solution of an inverse problem in magnetic resonance imaging using a regularized higher-order boundary element method, In: Boundary elements and other mesh reduction methods XXIX, Southampton: WIT Press, 2007, 323–332. https://doi.org/10.2495/BE070311 21. L. Marin, H. Power, R. W. Bowtell, C. Cobos-S´ anchez, A. A. Becker, P. Glover, et al., Boundary element method for an inverse problem in magnetic resonance imaging gradient coils, CMES Comput. Model. Eng. Sci.,23 (2008), 149–173. https://doi.org/10.3970/cmes.2008.023.149 AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1957 22. S. Moreno-Pulido, F. J. Garcia-Pacheco, C. Cobos-S´ anchez, A. Sanchez-Alzola, Exact solutions to the maxmin problem max kaxksubject to kbxk ≤ 1, Mathematics,8(2020), 85. http://doi.org/10.3390/math8010085 23. A. S´ anchez-Alzola, F. J. Garc´ ıa-Pacheco, E. Naranjo-Guerra, S. Moreno-Pulido, Supporting vectors for the `1-norm and the `∞-norm and an application, Math. Sci.,15 (2021), 173–187. https://doi.org/10.1007/s40096-021-00400-w 24. L. Surhone, M. Timpledon, S. Marseken, Principal component analysis: Karhunen-Lo`eve Theorem, Harold Hotelling, Karl Pearson, Exploratory Data Analysis, Eigendecomposition of a Matrix, Covariance Matrix, Singular Value Decomposition, Factor Analysis, Betascript Publishing, 2010. Supplementary: Python GUI code Let A∈Mm×n(R) be a general matrix. The algorithm to compute the supporting vector with its corresponding first principal component, the second eigenvector with its corresponding second principal component and the next eigenvectors and principal components of Ais the following: def PCA(X, num_components): X_meaned = X - np.mean(X, axis=0) cov_mat = np.cov(X_meaned, rowvar=False) eigen_values, eigen_vectors = np.linalg.eigh(cov_mat) sorted_index = np.argsort(eigen_values)[::-1] sorted_eigenvectors = eigen_vectors[:, sorted_index] eigenvector_subset = sorted_eigenvectors[:, 0:num_components] X_new_values = np.dot(eigenvector_subset.transpose(), X_meaned.transpose()).transpose() return X_new_values, eigenvector_subset One of the novelties of this work is to apply the PCA method with a different procedure specifically, using an algorithm based on the mathematical idea that the second eigenvector, with its associated second principal component, is the supporting vector of the original points projected to the orthogonal complement of the original supporting vector. Consequently, all the principal components can be computed in an iterative process via a supporting vector. Let x∈Rnbe a vector. First, we show the following function to calculate the orthogonal complement of x: def calculate_orthogonal_complement(x, normalize=True, threshold=1e-15): x = np.asarray(x) r, c = x.shape if r < c: import warnings warnings.warn(’fewer rows than columns’, UserWarning) s, v, d = np.linalg.svd(x) rank = (v > threshold).sum() AIMS Mathematics Volume 8, Issue 1, 1937–1958. 1958 oc = s[:, rank:] if normalize: k_oc = oc.shape[1] oc = oc.dot(np.linalg.inv(oc[:k_oc, :])) return oc Next, by using the previous function, we calculate the orthogonal complement of the supporting vector, representing a hyperplane with dimension n−1. Afterwards, the points of the original matrix, used for the initial supporting vector, are projected to this hyperplane. Hence, a new matrix is derived, whose first principal component corresponds to the second principal component of the original matrix. Thus, by applying this iterative algorithm, we can calculate all the principal components via the orthogonal complement of a supporting vector. It is easier to understand this algorithm considering a special case. Let A∈M3×3(R) be a matrix. In an intuitive manner, each row of the matrix represents points of R3. Now, we present a process of computing the first eigenvector (supporting vector) and its corresponding first principal component as before, but with a different procedure to obtain the followings. As mentioned, we notice that the first principal component via a supporting vector of a derived matrix from the original one coincides with the second principal component of the original matrix. This derived matrix is constructed projecting the original points of the original matrix to a plane formed by the orthogonal complement (two vectors) of the initial supporting vector (one vector) and the origin point. Therefore, we obtain the same result by applying the PCA presented before and this algorithm determining the next eigenvector as a supporting vector of the original supporting vector. supporting_vector = PCA(input_matrix, 1) orthogonal_complement = calculate_orthogonal_complement(supporting_vector) v1 = Vector(orthogonal_complement.T[0, :]) v2 = Vector(orthogonal_complement.T[1, :]) point = Point([0, 0, 0]) plane = Plane.from_vectors(point, v1, v2) projected_points = np.array([plane.project_point(x) for x in input_matrix]) supporting_vector_projected_points = PCA(projected_points, 1) ©2023 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0) AIMS Mathematics Volume 8, Issue 1, 1937–1958.