scieee AI-readable full text Open interactive document viewer

Development of hybrid ridge–PCA estimators for addressing Multicollinearity in Gaussian linear regression models

Alabi, Remilekun Enitan; Alabi, Olatayo Olusegun; Ojo, Oluwadare O

Abstract

This study tackles the persistent issue of multicollinearity in Gaussian linear regression which undermines the efficiency of Ordinary Least Squares (OLS) estimators. While Ridge Regression and Principal Component Analysis (PCA) are common remedies, they have limitations in terms of bias control and interpretability. To address this, the research proposes hybrid Ridge – PCA estimators using four newly developed ridge parameters combined with PCA. A Monte Carlo simulation evaluated 21 estimators including OLS, Ridge, PCA, and Liu estimators under varying sample sizes, error variances and multicollinearity levels using Mean Squared Error (MSE) as the performance metric. Results show that a newly hybrid estimator consistently outperformed other proposed and existing estimators by achieving the lowest MSE. The study demonstrates the strength of integrating regularization with dimensionality reduction to improve regression under multicollinearity.

Full text

 Corresponding author: Remilekun Enitan Alabi Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0. Development of hybrid ridge–PCA estimators for addressing Multicollinearity in Gaussian linear regression models Remilekun Enitan Alabi 1, *, Olatayo Olusegun Alabi 2 and Oluwadare O. Ojo 2 1 Department of Statistics, School of Science and Computer Studies, The Federal Polytechnic Ado-Ekiti, Ekiti state, Nigeria. 2 Department of Statistics, School of Physical Science, The Federal University of Technology, Akure, Ondostate, Nigeria. World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 Publication history: Received on 27 May 2025; revised on 05 July 2025; accepted on 07 July 2025 Article DOI: https://doi.org/10.30574/wjarr.2025.27.1.2559 Abstract This study tackles the persistent issue of multicollinearity in Gaussian linear regression which undermines the efficiency of Ordinary Least Squares (OLS) estimators. While Ridge Regression and Principal Component Analysis (PCA) are common remedies, they have limitations in terms of bias control and interpretability. To address this, the research proposes hybrid Ridge – PCA estimators using four newly developed ridge parameters combined with PCA. A Monte Carlo simulation evaluated 21 estimators including OLS, Ridge, PCA, and Liu estimators under varying sample sizes, error variances and multicollinearity levels using Mean Squared Error (MSE) as the performance metric. Results show that a newly hybrid estimator consistently outperformed other proposed and existing estimators by achieving the lowest MSE. The study demonstrates the strength of integrating regularization with dimensionality reduction to improve regression under multicollinearity. Keywords: Multicollinearity; Ridge Regression; Principal Component Estimator; Hybrid Estimators; Monte Carlo Simulation 1. Introduction The linear regression model is a statistical method that analyzes the relationship between an effect or dependent variable and one or more independent variables helping to explain or predict outcomes (Fayose et al., 2023b; Aladesuyi et al., 2025). The model is simply defined as follows: ,........ 22110 iikkiii xxxy  +++++= i = 1,…..,n, ……….. (1) where i y is the effect variable, 1i x ,…, ik x are the concomitant variables, 0  , 1  , …, k  are the unknown parameters to be estimated, i  denotes the disturbance term and it is assumed to be normally distributed with mean zero and constant variance . 2  The model is simply a simple regression model when there is one concomitant variable. The parameters in model (1) are mostly estimated by the Method of Least Squares (MLS). MLS is generally preferred and possesses some humbly properties when the assumptions of the linear regression models are intact, this makes the model to be classical (Owolabi et al., 2022 and Dawoud et al., 2022). These include linear relationship among the concomitant variables; the disturbance terms must come from a Gaussian distribution and has non scattered variance and others (Chatterjee and Hadi, 1977). In reality most of the aforementioned assumptions are normally violated. For instance, literature has shown that linear relationship often exists among concomitant variables which are termed multicollinearity (Fayose and Ayinde 2019; Shewa and Ugwuowo, 2022a). Multicollinearity is a phenomenon where World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 943 two or more concomitant variables are highly correlated in a Gaussian linear regression (Neter, et al. 1996; Fayose et al, 2023a and Aladesuyi et al., 2025). There is tendency for perfect or strong or moderate linear dependency among the concomitant variables (Shewa and Ugwuowo, 2022a; Shewa and Ugwuowo, 2022b). The method of least squares is unbiased but inefficient when there is linear relationship among the concomitant variables (Gujarati et al., 2012). It yields regression coefficients whose absolute values are too large and whose signs may actually reverse with negligible changes in the data (Buonaccorsi, 1996). If the multicollinearity is not perfect but high, the estimated coefficients can become unstable and highly sensitive to slight changes in the model leading to inflated standard errors and misleading inferences (Belseley et al., 1980; Fayose and Ayinde, 2019). Consequently, reliable interpretation of the model parameters may be compromised undermining the credibility of the results derived from such analysis. The pursuit of effective techniques to mitigate the adverse effects of multicollinearity is of paramount importance in both theoretical and applied statistics. Among the existing methods or approaches proposed to address or handle multicollinearity are Ridge Regression, Principal Component Analysis (PCA) among others. Both Ridge and PCA have emerged as prominent methodologies (Hoerl and Kennard, 1970; Jolliffe, 1986). Ridge regression offers a regularization technique that modifies the Ordinary least squares (OLSE) estimation process by introducing a penalty term (k) to the loss function thus allowing for the shrinkage of coefficient estimates towards zero (Hoerl and Kennard, 1970, Fayose et al., 2023b). The ridge parameter (k) counteracts the inflation of variances associated with multicollinearity effectively enhancing the stability of the estimates and producing more reliable and responsive predictions. In contrast, PCA serves as a dimensionality reduction technique that transforms the original correlated variables into a set of uncorrelated variables often called principal components (Jolliffe, 2002). By focusing on the principal components that explain the most variance in the data, PCA can help circumvent multicollinearity problem by ensuring that regression model utilizes orthogonal predictors. While both Ridge regression and PCA have demonstrated utility in addressing multicollinearity, each method presents notable limitations. Ridge regression while effective in controlling for multicollinearity does not completely eliminate correlation among predictors, it merely diminishes the variability of the coefficient estimates. Moreover, the choice of penalty parameter (k) can significantly influence the model’s performance necessitating careful cross – validation (Tikhonov, 1963). Meanwhile, PCA while adept at reducing dimensionality and addressing multicollinearity transforms the original predictors into a new set of components that may lack interpretability in the context of the original variables posing challenges for practical application and insight derivation (Jolliffe, 1986). To harness the strengths of both Ridge regression and PCA while mitigating their respective limitations, the concept of hybrid or combined estimators has been introduced by different authors in recent literature. These combined estimators integrate Ridge parameter (k) with PCA estimator to form or create a more robust framework or new hybrid estimator to tackle multicollinearity in Gaussian linear regression model. Wang et al., (2013) proposed a hybrid estimator that combines Ridge parameter (k) with PCA to improve parameter estimation in high dimensional settings. Their approach demonstrated a significant reduction in mean squared error (MSE) compared to standard methods. Other authors that have utilized these combined approaches are Zou and Hastie (2005), Buhlmann and Van de Geer (2011); Chang and Yang (2012); Zhang and Li (2015); Huang and Wang (2018) and Lukman, et al., (2020) among others. The paper intends to comprehensively propose new ridge parameter k’s and combine them with PCA to form new hybrid estimators to resolve the problem of multicollinearity within the Gaussian linear regression framework through simulation studies. We compared the estimators’ performance to that of some existing techniques. Section 2 contains the methodology. Section 3 presents the simulation design, while Section 4 presents simulation results and while Section 5 is the conclusion. 1.1. Existing and Proposed Estimators The matrix form of a linear regression model is defined as: eXy +=  ………………(2) Where y is the vector of response variables, X is 𝑛 × 𝑝 design matrix of concomitant variables or predictor 𝛽 is 𝑝×1 true vector or regression coefficients e ~ N ( ) 2 ,0  is the disturbance which is normally distributed with mean 0 and variance 2  . World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 944 1.2. The Ridge Estimator The Ordinary Least Square estimator is defined as: YXS I OLS 1 ˆ− =  ………… (3) where )( XXS I = ………… (4) The ridge regression estimator is the mostly used estimators in literature for handling multicollinearity problem. The Generalized Ridge Estimator is defined as: YXkISI GRRE 1 )( − +=  ………… (5) where S is a p x p product matrix of concomitant variables, 1 XY is a p x 1 vector of the product of effect and concomitant variables, k = diagonal ( 1 k , 2 k ,…, kp ), ki ≥ 0, i = 1, 2,..., p. k is a non – negative constant called biasing or ridge parameter. When k = 0, equation (5) returns to OLS estimator (Fayose and Ayinde, 2019; Kibria and Lukman, 2020; Fayose et al., 2023a). In this paper, we considered these selected ‘k’ parameters: Hoerl and Kennard (1975), Fayose and Ayinde (2019), Kibria and Lukman (2020), Chand and Kibria (2024a), Chand and Kibria (2024b) and also proposed four new ridge parameters respectively. 2. Principal Component Estimator The study considered PCA method to also handle multicollinearity in the model and also combined both ridge parameter and PCA method to form hybrid estimator to handle multicollinearity in the model. PCA transforms the original predictors x into a new set of uncorrelated variables called principal components. YXVSVVV PCR 11 )(  =−  ………… (6) ‘S’ is defined in equation (4). Let the covariance matrix of X be: XXC  = and the eigen value decomposition of C gives VVDC = ………… (7) where V is 𝑝 × 𝑝 matrix of eigen vectors (Principal Components) and D = diag ( ) p  .,.....,, 21 is the diagonal matrix of the eigen values. The data matrix X is transformed into principal components 𝑍 =𝑋𝑉. where Z is the new transformed data matrix of uncorrelated principal components By regressing y on the principal components Z instead of X; is given as: 𝑦=𝑍𝛾 +e ………… (8) Where 𝛾 = 𝑉′β The ordinary least square (OLS) estimator for 𝛾 in this model is given as; 𝛾 ( ) YZZZ  =−1 ………… (9) World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 945 By substituting Z = XV, therefore the PCA estimator of 𝛽 denoted as PCA  ˆ is defined as: PCA  ˆ = V𝛾 = ( ) yZZZV ' 1 '− ………… (10) Where ZZ = DXVXV = '' Therefore the PCA estimator in equation (10) can further be defined as: PCA  ˆ = ( ) yXVDV '' 1− ………… (11) 2.1. Some Alternative Ridge Estimators to OLSE The ridge estimator is defined as: ))( ˆ1yZkIZZ RIDGE  +  =−  ………… (12) where k is the non – negative tuning parameter. Different means of deriving k exists in the literature. These include: Following (Hoerl and Kennard, 1975), k is given by: ) ˆ ()( ˆ2 2 )( i M imedian medianHKkKGRHK   == , i = 1, 2, 3, p ………… (13) where 2 ˆ  = pn e n ii −  =1 2 is the mean square error from the MLS, i  is the ith element of the vector, and is the regression coefficient from the MLS. Following (Fayose and Ayinde, 2019), k is given by:                   −                 +         == 2 2 2 1 2 4 2 24 2 2 ˆ 2 ˆ ˆ ˆ 6 ˆ 4 ˆ ˆ ˆ )( ˆ         MiniMiniMini i Min iFAkKGRFA ………… (14) where .,...,3,2,1)( pMin iMin ==  Following (Kibria and Lukman, 2020), k is given by:                   + == i i Min iKLkKGRKL     ˆ ˆ 2 ˆ min)( ˆ 2 2 ………… (15) Following (Chand and Kibria, 2024), k is given by: )1( 11 ˆ )( ˆn p ipCKkKGRCK + ==  ………… (16) World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 946 Following (Chand and Kibria, 2024), k is given by: ),max( ˆ )( ˆ) 1 1( )1( 22 p n p ippCKkKGRCK + + ==  ………… (17) 2.1.1. The Liu Estimator The Liu estimator proposed by Liu (1993) combined the Stein estimator with Ordinary Ridge Regression estimator to handle multicollinearity. The Liu estimator of  is defined below as OLS III LdIXXIXX  ˆ )()( ˆ++= − 0 < d < 1. ………… (18) Where d =             +      2 2 2 ˆ ˆ ˆ min i i     ………… (19) Where d = ( )( i ddiag and is a diagonal matrix of the biasing parameter. The Liu estimator can return to OLS when d = 1. 2.2. Proposed Ridge Estimators For the ridge parameter whose estimators are defined in (13), (14), (15), (16) and (17), the concept of different forms by Lukman and Ayinde (2017) and Fayose and Ayinde (2019) was introduced based on minimum (MI), maximum (MA) and Median (MD) of eigen values (𝜆𝑖) of XIX of the design matrix of the regression model. Consequently, in this paper, we proposed some new ridge parameters whose estimators are defined below: RIDGE ESTIMATOR (PROPOSED 1) ),min( ˆ )2( ˆ) 1 1( )1( p n p ippCKk+ + =  ………… (20) i.e. The minimum version of Chand and Kibria (2024) RIDGE ESTIMATOR (PROPOSED 2)  =          + =p ii i HM KL pk PROP 12 2 2 ˆ )(     ………… (21) i.e. the Harmonic mean version of Kibria and Lukman (2020) RIDGE ESTIMATOR (PROPOSED 3)           + = i i Median KL mediank PROP     2 2 2 ˆ )( ………… (22) i.e. the median version of Kibria and Lukman (2020) RIDGE ESTIMATOR (PROPOSED 4 World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 947           + = )max( )max(2 ˆ2 2 )( i i FM KL PROP k     ………… (23) Fixed maximum of Kibria and Lukman (2020) 2.3. Derivation of the Properties of PCA with Ridge Estimator Ayinde et al. (2020) derived a new approach of Principal Component Analysis (PCA) estimator as an alternative to: PCA  ˆ = 𝑉𝐷−1𝑉′𝑋′𝑌 ………… (24) This is defined as PCA  ˆ = (𝑋′𝑋)−1𝑋′𝑦𝑟 ………… (25) Where r y ˆ is the predicted variable by regressing y on the r-principal component defined as ( ) yXZZZy rrrr  =−1 ˆ ………… (26) Such that r Z = r XT where Tr is the r – principal component and T is the orthogonal matrix. Therefore, combination of PCA with ridge estimator is defined as: 𝜷 𝑹−𝑷𝑪𝑨 =(𝑿′𝑿+𝒌𝜤)−𝟏𝑿′𝒁𝒓(𝒁𝒓′𝒁𝒓)−𝟏𝑿′𝒚 ………… (27) Where k is the biasing parameter for individual biasing parameter of Hoerl and Kennard (1975), Fayose and Ayinde (2019), Kibra and Lukman (2020), Chand and Kibra (2024a) and Chand and Kibra (2024b) 2.3.1. Properties of PCAR−  ˆ Mean of the PCAR−  ˆ To compute the mean of PCAR−  ˆ , the expected value of equation (3) is taking E(𝛽 󰆹𝑅−𝑃𝐶𝐴)=𝐸[(𝑋′𝑋+𝑘𝛪)−1𝑋′𝑍𝑟(𝑍𝑟′𝑍𝑟)−1𝑋′𝑦] ………… (28) eXy +=  =E( ( ) ( ) .) ˆ1 1       +         +  = − − −eXXZZZXkXXE rrrPCAR  = ( ) ( )                +  +         + − − − −eXZZZXkXXXXZZZXkXXE rrrrrr 1 1 1 1  ( ) ( ) )()( 1 1 1 1eEXZZZXkXXXEXZZZXkXX rrrrrr         +  +         +  = − − − −  World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 948 E(e) = 0 and  =)(E Therefore (𝐸(𝛽 󰆹𝑅−𝑃𝐶𝐴)=(𝑋′𝑋+𝑘𝛪)−1𝑋′𝑍𝑟(𝑍𝑟′𝑍𝑟)−1𝑋′.𝑋𝛽 ………… (29) Variance of the PCAR−  ˆ Var ( ) ˆPCAR−  = ( )( )       −− −−−− ) ˆ ( ˆ ) ˆ ( ˆPCARPCARPCARPCAR EEE  ………… (30) =𝐸[[(𝑋′𝑋+𝑘𝛪)−1𝑋′𝑍𝑟(𝑍𝑟′𝑍𝑟)−1𝑋′𝑦−(𝑋′𝑋+𝑘𝛪)−1𝑋′𝑍𝑟(𝑍𝑟′𝑍𝑟)−1𝑋′𝑋𝛽] [(𝑋′𝑋+𝑘𝛪)−1𝑋′𝑍𝑟(𝑍𝑟′𝑍𝑟)−1𝑋′𝑦−(𝑋′𝑋+𝑘𝛪)−1𝑋′𝑍𝑟(𝑍𝑟′𝑍𝑟)−1𝑋′𝑋𝛽]′.] ………… (31) Further expansion of (30) gives equation (31) ) ˆ (PCAR Var −  = 2  ( ) XXZZZXkXX rrr .)( 2 2 2        + − − ………… (32) Bias of the PCAR−  ˆ Bias( PCAR−  ˆ )= ( )  − −PCAR Eˆ ( ) . 1 1  −         +  = − −XXZZZXkXX rrr =((𝑋′𝑋+𝑘𝛪)−1𝑋′𝑍𝑟(𝑍𝑟′𝑍𝑟)−1𝑋′𝑋−𝛪)𝛽 ………… (33) Mean Squared Error Matrix of PCAR−  ˆ MSEM ( PCAR−  ˆ )= ( ) PCAR Var −  ˆ )+Bias2( PCAR−  ˆ ) ………… (34) MSEM( PCAR−  ˆ ) = 𝜎2(𝑋′𝑋+𝑘𝐼)−2(𝑋′𝑍𝑟)2(𝑍𝑟 ′𝑍𝑟)−1𝑋′𝑋+[(𝑋′𝑋+𝑘𝐼)−1𝑋′𝑍𝑟(𝑍𝑟 ′𝑍𝑟)−1𝑋′𝑋−𝐼]2𝛽2 ( ………… (35) Also equation (35) can still be written as: 𝑀𝑆𝐸𝑀(𝛽 󰆹𝑅−𝑃𝐶𝐴)=𝜎2(𝑋′𝑋+𝑘𝐼)−2(𝑋′𝑍𝑟)′(𝑋′𝑍𝑟)(𝑍𝑟 ′𝑍𝑟)−1𝑋′𝑋+[(𝑋′𝑋+𝑘𝐼)−1𝑋′𝑍𝑟(𝑍𝑟 ′𝑍𝑟)−1𝑋′𝑋−𝐼]′ [(𝑋′𝑋+𝑘𝐼)−1𝑋′𝑍𝑟(𝑍𝑟 ′𝑍𝑟)−1𝑋′𝑋−𝐼]𝛽′𝛽 ………… (36) Recall the general form of linear regression model in matrix form as given in (2) eXy +=  Therefore the canonical form of equation (2) can be written as: 𝑦=𝑊𝑋+𝑒 ………… (37) World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 949 Where  QXQW == , and Q is the orthogonal matrix whose columns indicate the eigen vectors of the design matrix .XX  Hence, ( ) p diagXQXQ  ,...,, 21 ==  where 0... 21  p  are the ordered eigen values .XX  The OLS estimator in equation (2) can be defined as: 𝛼𝑜𝑙𝑠 =∧−1 𝑋′𝑦 ………… (38) Thus the canonical form of equation (26) is ( ) yXZZZXk rrrPCAR         += − − − 1 1 ˆ  ………… (39) For the convenience of establishing the statistical properties of PCAR−  ˆ , the following Lemmas will be useful. Lemma I: Let F be a positive definite matrix such that F > 0 and let  be some vectors then; 0  −  F if and only if .1 1 −  F (Trenkler and Tontenbury,1990) Lemma II : Let ijjA=  ˆ for j=1,2 be two competing estimators of .  . Also suppose that ( ) ( ) 0 ˆˆ 21 −=  CovCovD where ( ) ( ) 21 ˆˆ  andCovCov are the covariance matrix of 21 ˆˆ  and Therefore ( ) ( ) 0 ˆˆ 21 −=  MSEMMSEMD if and only if   1 2 1 2 11 2 2+ −aaaDa  where ( ) 1 ˆ  MSEM = ( ) 1 ˆ  Cov + ' iiaa Such that i a =Bias ( ) ( )  IAiX i−= ˆ 2.3.2. The Superiority of the Proposed Estimator in the Sense of MSEM Criterion The proposed estimator is compared with some already existing estimators such as OLS, Ridge estimators in the sense of MSEM. Comparison between ols  ˆ and PCAR−  ˆ Recall the MSEM of the OLS estimator as ols  ˆ = yX  −1 as: MSEM( ols  ˆ )= 𝜎2𝛬−1 ………… (40) Equation (35) can be written as: MSEM( PCAR−  ˆ )=     1 ' 1111122 INNINNNNN rrrrrrr −−+ −−−−−− (41) where N= ( ) kI+ , r N = r ZX  , , 'rrr ZZ= XX = Therefore the difference between (39) and (40) is given as World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 950 MSEM( ols  ˆ )-MSEM( PCAR−  ˆ ) = PCAR ols D−   = 12 −   -      111 1 11122 INNINNNNN rrrrrrr −−+ −−−−−− = 12 −   -     INNINNNNN rrrrrrr −  −+ −−−−−− 11 ' 11122  =       INNINNNNN rrrrrrr −  −+  − −−−−−−− 11 1 111212  ………… (42) Where k>0 is the individual biasing parameter of Hoerl and Kennard (1970), Fayose and Ayinde (2019), Kibria and Lukman (2020) and Chand and Kibra (2024). The proposed estimator PCAR−  ˆ is superior to ols  ˆ if and only if       1 1 1 1 1 1212 ' 1 1−−− − − − − −− − −  INNNNNINN rrrrrrr Proof: By considering the dispersion matrix difference PCAR ols D−   =   − − −− 1 1212 rrr NNN  ………… (43) =trace ( PCAR ols D−   ) =  = p i diag 1 ( PCAR OLS D−   ) = 2   = p i diag 1   − − −− 1 121 rrr NNN = 2   = p i diag 1 ( ) p i iri ir ik n 1 2 2 1 =         + −    ………… (44) Where i  is the diag (X’X), nr =diag ( ) andZX r  ir  =diag ( ) r ZZ The difference will be positive definite if and only if ( ) irik  2 + - 2 2iir n  > 0. It can be observed that ( ) irik  2 + - 2 2iir n  will be greater than zero if k > 0. Hence by Lemma II the proof is completed Comparison between RE  ˆ and PCAR−  ˆ The bias vector covariance matrix and MSEM of RE  ˆ estimator defined as RE  ˆ = ( ) 1− + kI W’y ………… (45) World Journal of Advanced Research and Reviews, 2025, 27(01), 942-957 957 [31] Wang, Y., Li, R. and Zhang, H. H. (2013). A hybrid Ridge – PCA Estimator in High – Dimensonal Linear Models. Computational Statistics and Data Analysis, 57(1), 36 – 46. [32] Zhang, L. and Li, Y. (2015). Nonlinear Modeling with Principal Component Analysis and Penalized Regression Techniques. Journal of Computational and Graphical Statistics, 24(2), 372 – 392. [33] Zou, H. and Hastie, T. (2005). Regularization and Variable Selection via the Elastic Net. Journal of the Royal Statistical Society: Series B (Statistical Methodology). 67(2), 301 – 320.