A study on link functions for modelling and forecasting old-age survival probabilities of Australia and New Zealand
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Liu, Jacie Jia Article A study on link functions for modelling and forecasting old-age survival probabilities of Australia and New Zealand Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Liu, Jacie Jia (2021) : A study on link functions for modelling and forecasting oldage survival probabilities of Australia and New Zealand, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 9, Iss. 1, pp. 1-18, https://doi.org/10.3390/risks9010011 This Version is available at: https://hdl.handle.net/10419/258101 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
risks Article A Study on Link Functions for Modelling and Forecasting Old-Age Survival Probabilities of Australia and New Zealand Jacie Jia Liu Citation: Liu, Jacie Jia. 2021. A Study on Link Functions for Modelling and Forecasting Old-Age Survival Probabilities of Australia and New Zealand. Risks 9: 11. https://doi. org/10.3390/risks9010011 Received: 29 October 2020 Accepted: 29 December 2020 Published: 2 January 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). Department of Economics, Finance and Property, School of Business, Western Sydney University, Narellan Rd & Gilchrist Dr, Campbelltown, NSW 2560, Australia; [email protected] Abstract: Forecasting survival probabilities and life expectancies is an important exercise for actuaries, demographers, and social planners. In this paper, we examine extensively a number of link functions on survival probabilities and model the evolution of period survival curves of lives aged 60 over time for the elderly populations in Australasia. The link functions under examination include the newly proposed gevit and gevmin, which are compared against the traditional ones like probit, complementary log-log, and logit. We project the model parameters and so the survival probabilities into the future, from which life expectancies at old ages can be forecasted. We find that some of these models on survival probabilities, particularly those based on the new links, can provide superior fitting results and forecasting performances when compared to the more conventional approach of modelling mortality rates. Furthermore, we demonstrate how these survival probability models can be extended to incorporate extra explanatory variables such as macroeconomic factors in order to further improve the forecasting performance. Keywords: projection model; survival probability; life expectancy 1. Introduction The average human lifespan has been increasing consistently throughout the developed nations in the last hundred years or so. Oeppen and Vaupel (2002) reported that the highest female period life expectancy at birth around the world each year has been growing by about 0.24 years annually for more than a century. In Australia, period life expectancy at age 60 has grown from 16.5 years in 1950 to 25.0 years in 2017; in New Zealand, it has increased from 16.9 years in 1950 to 24.0 years in 2013. Life expectancy is one of the major indicators of a country’s wellbeing. Forecasting life expectancies accurately is of critical importance for government social planners as well as demographic researchers and industry practitioners (e.g., Lee et al. 1995;Nikolaevich 2019). The continuous mortality decline, together with the lack of clear improvement in the maximum lifespan, has caused the common phenomenon of a “rectangularisation” of the survival curve. This kind of observations has been well documented in earlier studies (e.g., Fries 1980;Cheung et al. 2005). As shown in Figure 1, it can be characterised by an upward and rightward shift of the period survival curve across successive time periods. The increasing area under the survival curve over time refers to the extent of rectangularisation. This concept is one of the major analytical frameworks in demographic research and can actually be further adapted to describe past changes in survival levels and predict future survival rates. The recent developments in mortality projection methods can provide a useful reference for this direction of exploiting the trends in survival probabilities. While there is a vast literature on modelling and projecting mortality rates (e.g., Lee and Carter 1992;Booth and Tickle 2008), relatively less attention has been paid on forecasting survival probabilities directly. Amongst the few previous works, De Jong and Marshall (2007) applied the probit link function to the survival probabilities of Australian females and assumed that it is driven by a single time trend. Hatzopoulos and Haberman (2015) used the complementary log-log link function and age-cohort effects within the Risks 2021,9, 11. https://doi.org/10.3390/risks9010011 https://www.mdpi.com/journal/risks
Risks 2021,9, 11 2 of 18 GLM (generalised linear model) framework for the cohort survival probabilities of a few European countries. Wong and Tsui (2015) proposed a new survival function for US women and men and modelled the changes of its parameters over time by autoregressive processes. Tan et al. (2016) constructed a hybrid survival curve and applied the logit link function to the annualised survival probabilities with two or three sets of time-varying parameters using Swedish and Bulgarian data. Risks 2021, 9, x FOR PEER REVIEW 2 of 20 females and assumed that it is driven by a single time trend. Hatzopoulos and Haberman (2015) used the complementary log-log link function and age-cohort effects within the GLM (generalised linear model) framework for the cohort survival probabilities of a few European countries. Wong and Tsui (2015) proposed a new survival function for US women and men and modelled the changes of its parameters over time by autoregressive processes. Tan et al. (2016) constructed a hybrid survival curve and applied the logit link function to the annualised survival probabilities with two or three sets of time-varying parameters using Swedish and Bulgarian data. There are potential advantages of modelling and projecting the survival probabilities rather than the mortality rates. First, when producing life expectancy forecasts, it would be more convenient to work with the survival probabilities directly without having to compound the mortality rates to form the survival rates. In a similar vein, when pricing pensions or annuities, the future probabilities of survival are a major input for the calculation process, and so it would be more natural to build a projection model that focuses on the survival probabilities. Moreover, as shown in the following sections, the survival probability patterns and trends are generally quite stable, which make their forecasting more straightforward than otherwise. We will show in our empirical analysis that some of these survival rate projection models can actually outperform the more conventional approach of a mortality rate projection model. As mentioned above, forecasting survival probabilities are much less explored in the literature when compared to forecasting mortality rates. In this paper, we attempt to reduce this knowledge gap by examining extensively a number of link functions on survival probabilities and modelling the evolution of the parameters and so the survival rates over time. The link functions considered include the probit, complementary log-log, logit, gevit functions, and a new link function based on the theory of minima, and both the age and period effects are incorporated. Furthermore, we illustrate how the survival rate projection models can be extended to include additional explanatory variables such as the GDP (gross domestic product) per capita. Some previous examples of incorporating macroeconomic factors into mortality rate projection models are Hanewald (2011); Niu and Melenberg (2014); French and O’Hare (2014); and Seklecka et al. (2019). To the best of our knowledge, our paper provides the first attempt to incorporate macroeconomic factors into the survival rate projection models. The remainder of the papers is as follows. Section 2 gives an introduction of the various survival rate projection models being considered. Section 3 compares the fitting performances of the models for Australian and New Zealand data. Section 4 sets forth an outof-sample analysis for measuring the forecasting performances of the models. Section 5 studies the potential link between the survival probabilities and the economic growth. Section 6 concludes. Figure 1. Survival curves of New Zealanders aged 60 in 1970, 1980, 1990, 2000, and 2013. There are potential advantages of modelling and projecting the survival probabilities rather than the mortality rates. First, when producing life expectancy forecasts, it would be more convenient to work with the survival probabilities directly without having to compound the mortality rates to form the survival rates. In a similar vein, when pricing pensions or annuities, the future probabilities of survival are a major input for the calculation process, and so it would be more natural to build a projection model that focuses on the survival probabilities. Moreover, as shown in the following sections, the survival probability patterns and trends are generally quite stable, which make their forecasting more straightforward than otherwise. We will show in our empirical analysis that some of these survival rate projection models can actually outperform the more conventional approach of a mortality rate projection model. As mentioned above, forecasting survival probabilities are much less explored in the literature when compared to forecasting mortality rates. In this paper, we attempt to reduce this knowledge gap by examining extensively a number of link functions on survival probabilities and modelling the evolution of the parameters and so the survival rates over time. The link functions considered include the probit, complementary log-log, logit, gevit functions, and a new link function based on the theory of minima, and both the age and period effects are incorporated. Furthermore, we illustrate how the survival rate projection models can be extended to include additional explanatory variables such as the GDP (gross domestic product) per capita. Some previous examples of incorporating macroeconomic factors into mortality rate projection models are Hanewald (2011); Niu and Melenberg (2014); French and O’Hare (2014); and Seklecka et al. (2019). To the best of our knowledge, our paper provides the first attempt to incorporate macroeconomic factors into the survival rate projection models. The remainder of the papers is as follows. Section 2gives an introduction of the various survival rate projection models being considered. Section 3compares the fitting performances of the models for Australian and New Zealand data. Section 4sets forth an out-of-sample analysis for measuring the forecasting performances of the models. Section 5 studies the potential link between the survival probabilities and the economic growth. Section 6concludes. 2. Survival Rate Projection Models Suppose qx,t is the mortality rate at age x in year t , and npx,t=∏n i=1(1−qx+i−1,t) is the corresponding period survival probability in year t for a surviving period of at least
Risks 2021,9, 11 3 of 18 n years. The first link function we consider is the probit function used by De Jong and Marshall (2007). We set the model structure as Φ−1(npx0,t) = h(x , t) , in which Φ−1 is the inverse standard normal cdf (cumulative distribution function) and h(x , t) is a regression structure allowing for the age and period effects with x=x0+n . If one treats 1 −npx0,t as the cdf of the future lifetime (within n years) of a life aged x0 in year t , npx0,t as the survival function of that future lifetime, and h(x , t2)−h(x , t1) = λ as a certain constant for t2>t1 , it can be deduced that npx0,t2=Φ(Φ−1(npx0,t2))= Φ(h(x , t2)) = Φ(h(x , t1) + λ)= Φ(Φ−1(npx0,t1) + λ) and so 1 −npx0,t2=Φ(Φ−1( 1 −npx0,t1)−λ) . It then means that under this probit model structure, the future death rates can be seen as a Wang transform (Wang 2000) of the past death rates, with the parameter λ capturing the mortality decline. Note that the probit link function ensures that the survival rates are between 0 and 1, regardless of the values of h(x , t) , and that it is a symmetric link as Φ(z) approaches 0 and 1 at the same pace. The next one is the complementary log-log link function used by Hatzopoulos and Haberman (2015). We follow their model structure as ln(−ln( 1 −( 1 −npx0,t))) = ln(−ln(npx0,t)) = h(x , t) , that is, the link function is being applied on 1 −npx0,t , but not npx0,t . Unlike the probit function, the complementary log-log function is an asymmetric link, which would be useful for the rectangularisation patterns as in Figure 1. This asymmetry may suit survival modelling better when compared to a symmetric one. The survival rates under this link function are between 0 and 1. We apply the logit link function from Tan et al. (2016) as ln(npx0,t/( 1 −npx0,t)) = h(x , t) . It is a symmetric link like the probit function, and it constrains the survival rates between 0 and 1. Its inverse function exp(z)/( 1 +exp(z)) is actually the cdf of the logistic distribution with the location parameter equal to 0 and the scale parameter equal to 1. Both the probit and logit functions are used extensively in binary response models. Recently, Medford and Vaupel (2019) proposed the so-called gevit link function for modelling mortality rates. We apply this link function differently here as [(−ln(npx0,t))−ξ− 1 ]/ξ=h(x , t) ; that is, it is being applied on the survival probability but not on the death rate. This link is asymmetric and it constrains the survival rates like those above. Its inverse function exp(−(1+ξz)−1/ξ) is indeed the cdf of the GEV (generalised extreme value) distribution for maxima with the location parameter equal to 0 and the scale parameter equal to 1. The extra shape parameter ξ provides more flexibility to manage the extent of asymmetry. Inspired by the gevit link function, which is based on maxima, we exploit the GEV distribution for “minima” instead (e.g., Liu and Li 2019) as an alternative. The cdf of the minima GEV distribution is specified as 1 −exp(−(1−ξz)−1/ξ) . Accordingly, we consider a new model structure [ 1 −(−ln(1−npx0,t))−ξ]/ξ=h(x , t) . It is straightforward to deduce that this new link is also asymmetric and the resulting survival rates are within the valid range. We refer to it as the “gevmin” model structure in the following analysis. If this new link is applied on 1 −npx0,t instead of npx0,t , the model structure becomes [ 1 −(−ln(1−(1−npx,t)))−ξ]/ξ= [ 1 −(−ln(npx0,t))−ξ]/ξ=h(x , t) . Then if ξ= 0, it reduces to ln(−ln(npx,t)) = h(x , t) , which means that the complementary log-log model structure above can effectively be seen as a specific example of this alternative model structure from minima. Figure 2compares the symmetry of the probit and logit functions with the asymmetry of the complementary log-log, gevit, and gevmin functions as described above. For the symmetric ones, the response approaches 0 and 1 at the same pace. For the asymmetric complementary log-log and gevit functions, the response approaches 1 slower than reaching 0. By contrast, under the new gevmin model structure, the response of the (inverse) link approaches 1 faster than reaching 0, which is an opposite situation. As reflected in Figure 1, as the rectangularisation continues to occur, the survival rates of more and more ages x rise above 0.5 and get closer to 1 over time t , while the rates drop more and more sharply at the progressively narrowing highest end. This phenomenon may make one or more of the asymmetric candidates more suitable for modelling how the survival rates evolve
Risks 2021,9, 11 4 of 18 over time. More details of the maxima and minima GEV distributions are given in the Appendix A. Risks 2021, 9, x FOR PEER REVIEW 5 of 20 Figure 2. Five different (inverse) link functions. Figure 2. Five different (inverse) link functions.
Risks 2021,9, 11 5 of 18 In addition to npx0,t , we also follow Tan et al. (2016) and consider another response (npx0,t)1 n , that is, the “annualised” survival probability for comparison. Figure 3shows that the resulting annualised survive curve also displays an upward and rightward shift over successive time periods. Regarding the regression structure h(x , t) , we employ both the Lee and Carter (1992) structure and the Cairns et al. (2009) structure to allow for the age and period effects. The first one is taken as h(x,t) = ax+bxkt, and the second one as h(x , t) = kt,1 +kt,2(x−x)+kt,3((x−x)2−σ2) , where ax is the age effect, kt is the “survival index” with age-specific sensitivity bx , kt,1 to kt,3 are three time-varying parameters, x is the average of the age range, and σ2 is the average of (x−x) over the age range considered. Altogether, there are 5 (links) × 2 (responses) × 2 (structures) = 20 combinations under our consideration. This coverage is much more comprehensive than those of the few earlier papers on projecting the survival rates directly. Subsequently, we will also explore adding macroeconomic factors into the regression structure to see whether it can improve the performances. The Appendix Aprovides the parameter estimation methods for the models tested. Risks 2021, 9, x FOR PEER REVIEW 6 of 20 Females Males Figure 3. Annualised survival curves of Australians aged 60 in 1970, 1980, 1990, 2000, and 2017. 3. Fitting Performances In this section, we apply the 20 models (combinations) to female and male populations in Australia and New Zealand. The mortality data of ages 60 to 99 and years 1970 to 2017 of Australia (1970 to 2013 for New Zealand) are extracted from the Human Mortality Database (HMD 2020). For demonstration purposes, we consider the survival probabilities for a life aged 60 and set 060x= (for surviving periods n = 1, 2, 3, …, 40), as the mortality rates below age 60 have already reached very low levels and have relatively little impact on the overall survival rates. In fact, the survival rate of a newborn for a surviving period up to 60 years is now very close to one, and the corresponding movements over time look too insignificant for a meaningful projection. By contrast, there is still much room for the survival rates at ages 60 and above to improve, and as shown below, the resulting patterns and trends are rather stable and can readily be projected into the future. Moreover, longevity of retirees is a serious concern for governments, insurers selling annuities, and pension funds because of the increasing financial burden. It would be of high practical interest to focus on the mortality of retirement ages. Figure 4 plots the survival index t k of the first regression structure and also the time-varying parameters ,1t k to ,3t k of the second regression structure based on the five different link functions for the survival probabilities of Australian females. It is interesting to see that all the temporal parameters of t k and ,1t k demonstrate a strong linearly increasing trend. It reflects clearly the continually improving survival rates at old ages as a whole. (For the complementary log-log model structure, Figure 2 displays a negative relationship between the response and the argument in contrast to the others. The resulting major trends of t k and ,1t k in Figure 4 are then inverted, but the implication on mortality decline is the same.) The time-varying parameters ,2t k and ,3t k of the second regression structure refer to the slope and curvature in year t. While the directions of their trends are different because of the differences in how the link functions operate on the survival rates, a high level of linearity can largely be observed. We can then use the (univariate or multivariate) random walk with drift to project all these linear trends. Compared to the time-varying trends usually seen when modelling the mortality rates (e.g., Cairns et al. 2009), modelling the survival rates here generates more linear trends, which make the use of the random walk with drift more justifiable than otherwise. The observations for Australian male and New Zealand populations (not shown here) exhibit similar patterns. Figure 3. Annualised survival curves of Australians aged 60 in 1970, 1980, 1990, 2000, and 2017. Note that some of these stochastic projection models can be seen as modifications of those standard static mortality or survival functions. Suppose µx and tpx represent the force of mortality and the t-year survival rate of a life aged xfor a static one-year life table. For instance, the Gompertz law states that µx=αβx , which can be transformed into the survival function tpx=exp(−αβx(βt− 1 )/ ln β) and a regression structure ln µx=ln α+xln β . The old-age component of the Heligman–Pollard curve can be taken as qx/px=αβx , which can be expressed as the survival function px= (1+αβx)−1 and a regression structure logit qx=ln α+xln β . The famous Lee–Carter model (Lee and Carter 1992) and the CBD model (Cairns et al. 2006) can be seen in some way as adding time-varying components into the two regression structures above (log and logit) in order to turn them from being static to stochastic. While the stochastic Lee–Carter and CBD models deal with mortality rates, as compared to our approach of modelling survival probabilities, we will include them in the following analysis for comparison. 3. Fitting Performances In this section, we apply the 20 models (combinations) to female and male populations in Australia and New Zealand. The mortality data of ages 60 to 99 and years 1970 to 2017 of Australia (1970 to 2013 for New Zealand) are extracted from the Human Mortality Database (HMD 2020). For demonstration purposes, we consider the survival probabilities for a life aged 60 and set x0= 60 (for surviving periods n = 1, 2, 3, . . . , 40), as the mortality rates below age 60 have already reached very low levels and have relatively little impact on
Risks 2021,9, 11 6 of 18 the overall survival rates. In fact, the survival rate of a newborn for a surviving period up to 60 years is now very close to one, and the corresponding movements over time look too insignificant for a meaningful projection. By contrast, there is still much room for the survival rates at ages 60 and above to improve, and as shown below, the resulting patterns and trends are rather stable and can readily be projected into the future. Moreover, longevity of retirees is a serious concern for governments, insurers selling annuities, and pension funds because of the increasing financial burden. It would be of high practical interest to focus on the mortality of retirement ages. Figure 4plots the survival index kt of the first regression structure and also the timevarying parameters kt,1 to kt,3 of the second regression structure based on the five different link functions for the survival probabilities of Australian females. It is interesting to see that all the temporal parameters of kt and kt,1 demonstrate a strong linearly increasing trend. It reflects clearly the continually improving survival rates at old ages as a whole. (For the complementary log-log model structure, Figure 2displays a negative relationship between the response and the argument in contrast to the others. The resulting major trends of kt and kt,1 in Figure 4are then inverted, but the implication on mortality decline is the same.) The time-varying parameters kt,2 and kt,3 of the second regression structure refer to the slope and curvature in year t . While the directions of their trends are different because of the differences in how the link functions operate on the survival rates, a high level of linearity can largely be observed. We can then use the (univariate or multivariate) random walk with drift to project all these linear trends. Compared to the time-varying trends usually seen when modelling the mortality rates (e.g., Cairns et al. 2009), modelling the survival rates here generates more linear trends, which make the use of the random walk with drift more justifiable than otherwise. The observations for Australian male and New Zealand populations (not shown here) exhibit similar patterns. Table 1reports the MAPE (mean absolute percentage error) values on np60,t of fitting the 20 models, with the original Lee–Carter model (Lee and Carter 1992) and the CBD model with curvature (Cairns et al. 2009) for the death rates also included for comparison. The major observations are stated below: (1) For the first regression structure, the MAPE values are often smaller when the response is npx0,t . By contrast, for the second regression structure, the MAPEs are much smaller when the response is (npx0,t)1 n. (2) When the response is npx0,t , the MAPE values from the first regression structure are clearly smaller. However, when the response is (npx0,t)1 n , the situation is mostly reversed. (3) For the complementary log-log model structure with the first regression form, the MAPE remains the same regardless of the response (i.e., npx0,t or (npx0,t)1 n ). The underlying reason is that ln(−ln((npx0,t)1 n)) = ax+bxkt is indeed equivalent to ln(−ln(npx0,t))= ax+bxkt+ln n=a∗ x+bxkt , where a∗ x=ax+ln n . Hence, they produce the same fitted values of np60,tand so the same MAPEs. (4) Overall, the combination of the response (npx0,t)1 n and the second regression structure (i.e., kt,1 +kt,2(x−x)+kt,3((x−x)2−σ2) ) gives the better MAPEs. In particular, the gevmin model structure, based on the newly proposed gevmin link, leads to the smallest MAPEs consistently for all the populations (0.73, 0.93, 1.20, 1.68) considered. (5) The best gevmin model structures noted above outperform the more traditional approaches of the Lee–Carter and CBD models in modelling the mortality rates (with MAPEs of 1.39 to 3.02).
Risks 2021,9, 11 7 of 18 Risks 2021, 9, x FOR PEER REVIEW 8 of 20 to 1989 (20 years), 1970 to 1994 (25 years), 1970 to 1999 (30 years), and 1970 to 2004 (35 years) and then forecast the survival rates for the remaining periods until the very last year of available data. Here we focus on the annualised response and the second regression structure since they give the better fitting performances in Table 1. In addition to continuing to use the Lee–Carter and CBD models for comparison, we also apply the multivariate random walk with drift to the survival probabilities as a benchmark, naive model. The MAPEs of not only the projected 60,nt p but also the projected life expectancies at age 60 are examined to compare the performances. They are calculated as 1 60, 60, 60, ,ˆ ||/ d ntntnt nxt ppp− and 1 60, 60, 60, ˆ ||/ tttt nt eee− respectively, where 60, ˆ nt p and 60,nt p ( 60, ˆ t e and 60,t e ) are the projected and observed survival probabilities (life expectancies) in year t , d n is the number of data points, and t n is the number of years in the testing period. For each case (column) in Table 2, the three lowest MAPE values are highlighted (in grey). We notice the following patterns within the results in Table 2: (1) Out of the 16 cases (4 fitting periods × 2 countries × 2 sexes), the gevit and gevmin model structures produce the three lowest MAPEs in 12 cases. Their performances are the most consistent ones amongst all the candidates. (2) For the gevit and gevmin model structures, the average MAPE is 6.10. For the probit, complementary log-log, and logit model structures, the average MAPE is 6.29. For the LC and CBD models, the average MAPE is 6.53. For the naive random walk model, the average MAPE is 7.95. (3) The MAPE values tend to be lower for females (4.04 on average) than for males (8.98 on average). (4) The MAPE values tend to be lower for Australia (5.65 on average) than for New Zealand (7.36 on average). (5) The naive random walk model leads to the highest MAPE values in 10 cases. probit model structure complementary log-log model structure Risks 2021, 9, x FOR PEER REVIEW 9 of 20 logit model structure gevit model structure gevmin model structure Figure 4. Trends of t k and ,1t k to ,3t k based on five link functions for non-annualised survival probabilities of Australian females. Table 1. MAPE values (%) on fitted survival probabilities at age 60 from fitting 22 models to Australian and New Zealand data (1970–2017 and 1970–2013 respectively). Model Australian Females Australian Males Non-Annualised Annualised Non-Annualised Annualised probit − t k 1.24 1.47 2.03 2.53 log-log − t k 1.59 1.59 2.75 2.75 logit − t k 1.39 1.59 1.94 2.72 gevit − t k 1.11 1.09 1.91 2.00 gevmin − t k 1.12 1.14 1.82 1.87 probit − ,1t k − ,3t k 4.38 1.14 5.76 1.12 log-log − ,1t k − ,3t k 7.20 2.17 11.74 1.80 logit − ,1t k − ,3t k 8.49 1.96 11.76 1.60 gevit − ,1t k − ,3t k 2.71 0.85 3.48 1.01 gevmin − ,1t k − ,3t k 2.95 0.73 3.71 0.93 Figure 4. Trends of kt and kt,1 to kt,3 based on five link functions for non-annualised survival probabilities of Australian females.
Risks 2021,9, 11 8 of 18 Table 1. MAPE values (%) on fitted survival probabilities at age 60 from fitting 22 models to Australian and New Zealand data (1970–2017 and 1970–2013 respectively). Model Australian Females Australian Males NonAnnualised Annualised NonAnnualised Annualised probit −kt1.24 1.47 2.03 2.53 log - log −kt1.59 1.59 2.75 2.75 logit −kt1.39 1.59 1.94 2.72 gevit −kt1.11 1.09 1.91 2.00 gevmin −kt1.12 1.14 1.82 1.87 probit −kt,1−kt,3 4.38 1.14 5.76 1.12 log - log −kt,1−kt,3 7.20 2.17 11.74 1.80 logit −kt,1−kt,3 8.49 1.96 11.76 1.60 gevit −kt,1−kt,3 2.71 0.85 3.48 1.01 gevmin −kt,1−kt,3 2.95 0.73 3.71 0.93 Lee-Carter 1.39 2.42 CBD 1.62 1.57 Model New Zealand Females New Zealand Males NonAnnualised Annualised NonAnnualised Annualised probit −kt2.23 2.61 2.76 3.11 log - log −kt2.82 2.82 3.27 3.27 logit −kt2.36 2.81 2.77 3.24 gevit −kt2.06 2.02 2.70 2.73 gevmin −kt2.09 2.11 2.66 2.70 probit −kt,1−kt,3 4.80 1.45 5.95 1.86 log - log −kt,1−kt,3 8.21 2.24 12.65 2.35 logit −kt,1−kt,3 9.22 2.06 12.19 2.20 gevit −kt,1−kt,3 3.09 1.26 3.53 1.71 gevmin −kt,1−kt,3 3.38 1.20 3.71 1.68 Lee-Carter 2.44 3.02 CBD 2.00 2.15 Compared with the traditional link functions, the additional shape parameter ξ in the gevmin and gevit link functions makes them a lot more flexible in capturing different degrees of asymmetry. For each population in Table 1, the two smallest MAPEs are all generated from either the gevmin or gevit model structure. The empirical advantage of using these two newer links is apparent here. Note also that the response (npx0,t)1 n (see Figure 3again) has a simpler shape over n than the response npx0,t (see Figure 1). Hence, a simple regression structure in terms of merely x (i.e., kt,1 +kt,2(x−x)+kt,3((x−x)2− σ2) ) would suffice for the former, while a more dedicated allowance for the age effect (i.e., ax+bxkt) would be needed for the latter. As noted earlier, if the new gevmin link is applied on 1 −npx0,t rather than npx0,t , the model structure turns into [ 1 −(−ln(npx0,t))−ξ]/ξ=h(x , t) . It can further be rearranged as [(−ln(npx0,t))−ξ− 1 ]/ξ=−h(x , t) = h∗(x , t) , which then becomes the gevit model structure effectively. Consequently, they generate the same fitted values of np60,t and the same MAPEs, though the resulting signs and trends of their parameters are opposite to each other. However, the increasing survival index based on the gevmin link on npx0,t has a more natural interpretation in terms of improving survival over time. 4. Forecasting Performances In this section, we perform an out-of-sample analysis to assess the forecast accuracy of the survival rate projection models. We apply the models to four fitting periods of 1970 to 1989 (20 years), 1970 to 1994 (25 years), 1970 to 1999 (30 years), and 1970 to 2004 (35 years) and then forecast the survival rates for the remaining periods until the very
Risks 2021,9, 11 15 of 18 Second, while we have used two particular regression forms in our analysis, there are other possible structures that also allow for the age and period effects. Some examples include adding extra time-varying parameters and co-modelling the survival rates of related subpopulations. Funding: This research received no external funding. Institutional Review Board Statement: Not applicable. Informed Consent Statement: Not applicable. Data Availability Statement: Not applicable. Conflicts of Interest: The author declares no conflict of interest. Appendix A Suppose X1 , X2 , . . . , Xn are n independent and identically distributed (i.i.d.) random variables having cdf F(x) . Let Mn=max(X1 , X2 , . . . , Xn) be the maximum value amongst them. Under the Fisher–Tippett–Gnedenko Theorem and certain technical conditions, when n goes to infinity, (Mn−bn)/an (for some an> 0 and real bn ) follows the so-called GEV(µ , σ , ξ) distribution (for −∞<µ<∞ and σ> 0). There are three specific cases: Gumbel ( ξ= 0), Fréchet ( ξ> 0), and reversed Weibull ( ξ< 0). Their cdfs G(x) are, respectively, G(x) = exp(−exp(−x−µ σ)),(for ξ=0 , −∞<x<∞) G(x) = exp(−(1+ξx−µ σ) −1 ξ),(for ξ>0, x>µ−σ ξ) G(x) = exp(−(1+ξx−µ σ) −1 ξ).(for ξ<0, x<µ−σ ξ) Let Un=min(X1 , X2 , . . . , Xn) be the corresponding minimum value. Under some technical conditions, with n approaching infinity, (Un−bn)/an follows the GEVmin(µ , σ , ξ) distribution (for −∞<µ<∞ and σ> 0). There are also three specific cases: reversed Gumbel ( ξ= 0), reversed Fréchet ( ξ> 0), and Weibull ( ξ< 0). Their cdfs H(x) are, respectively, H(x) = 1−exp(−exp(x−µ σ)),(for ξ=0 , −∞<x<∞) H(x) = 1−exp(−(1−ξx−µ σ) −1 ξ),(for ξ>0, x<µ+σ ξ) H(x) = 1−exp(−(1−ξx−µ σ) −1 ξ).(for ξ<0, x>µ+σ ξ) The SVD (singular value decomposition) method is employed to estimate the parameters in h(x , t) = ax+bxkt . Regarding h(x , t) = kt,1 +kt,2(x−x)+kt,3((x−x)2−σ2) , the least squares method is used to estimate the parameters for each year t . When there exists a shape parameter ξ , an extensive range of different trial values are taken for it in turn and then different sets of parameter values are estimated accordingly. The final value of the shape parameter is determined by the one giving the smallest overall fitting error. Table A1 below provides the final values of the shape parameter for different populations and fitting periods. All the estimated values are negative, implying that the GEV distribution involved belongs to the reversed Weibull. For the second regression structure (right value), all the estimated values are quite consistent with one another, ranging from about − 0.4 to − 0.3. For the first regression structure (left value), the estimated values range from − 0.6 to − 0.5 when the response is npx0,t and are − 1.1 to − 0.8 when the response is (npx0,t)1 n . Note that
Risks 2021,9, 11 16 of 18 these estimation approaches are adopted in this paper because each response depends on the mortality rates of n ages, so it is impossible to apply the Poisson assumption via maximum likelihood that is commonly used with mortality rate projection models. Table A1. Estimated shape parameter values of gevit model structure fitted to Australian and New Zealand data (left value—first regression structure; right value—second regression structure). Fitting Period Australian Females Australian Males Non-Annualised Annualised Non-Annualised Annualised 1970–2017 −0.59/−0.32 −0.85/−0.31 −0.60/−0.39 −0.93/−0.30 1970–1999 −0.66/−0.35 −0.99/−0.42 −0.62/−0.45 −1.08/−0.40 1970–1989 −0.57/−0.34 −0.89/−0.37 −0.62/−0.43 −1.06/−0.37 Fitting Period New Zealand Females New Zealand Males Non-Annualised Annualised Non-Annualised Annualised 1970–2013 −0.58/−0.34 −0.90/−0.31 −0.53/−0.41 −0.86/−0.35 1970–1999 −0.54/−0.37 −0.83/−0.37 −0.48/−0.45 −0.86/−0.45 1970–1989 −0.55/−0.36 −0.86/−0.36 −0.53/−0.43 −0.91/−0.38 As noted in Hyndman and Koehler (2006), the MAPE would have the disadvantage of putting a heavier penalty on positive errors than on negative errors. Accordingly, the sMAPE values are also calculated and reported in Tables 2and 3below. Regarding the sMAPEs of the projected survival probabilities in Table 2, the average sMAPE is 6.57 for the gevit and gevmin model structures, compared to 7.01 for the probit, complementary log-log, and logit model structures, 6.93 for the LC and CBD models, and 9.23 for the naive random walk model. Regarding the sMAPEs of the projected life expectancies in Table 3 , the average sMAPE is 2.35 for the gevit and gevmin model structures, compared to 2.62 for the probit, complementary log-log, and logit model structures, 2.89 for the LC and CBD models, and 2.50 for the naive random walk model. The gevit and gevmin model structures contribute to the three lowest sMAPEs in 11 or 12 of the 16 cases, and overall, they show the best forecasting performances in our out-of-sample analysis on the survival rates. Table 2. sMAPE values (%) on projected survival probabilities at age 60 using Australian and New Zealand data for different fitting periods. Australian Females Australian Males Model 1970–1989 1970–1994 1970–1999 1970–2004 1970–1989 1970–1994 1970–1999 1970–2004 probit 5.79 1.87 2.89 2.90 17.67 11.25 2.76 1.67 log-log 6.91 2.13 2.60 1.98 18.82 12.39 3.64 1.63 logit 6.71 2.03 2.58 2.70 18.59 12.16 3.42 2.06 gevit 4.75 2.52 3.91 3.68 16.27 10.01 2.41 2.04 gevmin 4.75 2.35 3.67 3.49 16.73 10.40 2.41 1.74 LC 3.92 3.08 3.00 2.29 11.72 8.07 4.13 3.33 CBD 4.90 1.89 3.10 3.32 18.49 12.58 3.82 2.34 MRW 9.88 4.72 2.52 2.03 23.17 17.20 8.55 5.87 New Zealand Females New Zealand Males Model 1970–1989 1970–1994 1970–1999 1970–2004 1970–1989 1970–1994 1970–1999 1970–2004 probit 5.06 5.21 3.41 2.90 15.20 9.25 10.79 9.36 log-log 5.56 4.76 3.68 3.49 15.75 9.79 11.41 9.90 logit 5.44 4.84 3.60 3.35 15.63 9.68 11.27 9.78 gevit 4.97 5.65 3.55 2.95 14.36 8.57 10.16 8.97 gevmin 4.98 5.67 3.55 2.93 14.64 8.83 10.35 9.04 LC 6.37 6.21 4.92 3.35 18.57 10.26 9.43 6.82 CBD 5.64 7.00 3.78 2.91 16.41 9.92 11.40 8.73 MRW 7.19 3.15 3.80 3.57 18.79 12.51 13.26 11.39
Risks 2021,9, 11 17 of 18 Table 3. sMAPE values (%) on projected life expectancies at age 60 using Australian and New Zealand data for different fitting periods. Australian Females Australian Males Model 1970–1989 1970–1994 1970–1999 1970–2004 1970–1989 1970–1994 1970–1999 1970–2004 probit 1.69 0.69 1.11 1.19 5.51 3.75 0.84 0.55 log-log 2.02 0.71 0.87 1.02 6.10 4.33 1.35 0.39 logit 1.99 0.71 0.89 1.04 6.03 4.27 1.29 0.40 gevit 1.25 0.95 1.59 1.50 4.44 2.74 0.66 1.02 gevmin 1.35 0.81 1.41 1.37 4.98 3.28 0.66 0.77 LC 2.07 0.74 0.79 0.81 6.33 4.22 1.56 0.70 CBD 2.25 0.79 0.78 0.99 6.63 4.63 1.70 0.52 MRW 1.34 0.79 1.41 1.47 5.20 3.47 0.74 0.68 New Zealand Females New Zealand Males Model 1970–1989 1970–1994 1970–1999 1970–2004 1970–1989 1970–1994 1970–1999 1970–2004 probit 2.88 0.75 1.80 1.18 7.49 4.41 4.19 2.22 log-log 2.98 0.81 1.85 1.15 7.79 4.71 4.42 2.36 logit 2.97 0.80 1.85 1.15 7.75 4.67 4.40 2.35 gevit 2.67 0.78 1.65 1.16 6.89 3.88 3.78 1.93 gevmin 2.74 0.75 1.72 1.19 7.16 4.14 3.98 2.10 LC 3.57 0.92 2.41 1.13 9.87 5.46 4.64 2.27 CBD 3.18 0.80 1.80 1.17 7.70 4.93 4.47 2.57 MRW 3.19 0.80 1.76 1.15 7.51 4.37 4.06 2.03 References Boonen, Tim J., and Hong Li. 2017. Modeling and forecasting mortality with economic growth: A multipopulation approach. Demography 54: 1921–46. [CrossRef] [PubMed] Booth, Heather, and Leonie Tickle. 2008. Mortality modelling and forecasting: A review of methods. Annals of Actuarial Science 3: 3–43. [CrossRef] Box, George E. P., and Gwilym M. Jenkins. 1976. Time Series Analysis: Forecasting and Control, 2nd ed. San Francisco: Holden-Day. Brenner, Harvey M. 2005. Commentary: Economic growth is the basis of mortality rate decline in the 20th century: Experience of the United States 1901–2000. International Journal of Epidemiology 34: 1214–21. [CrossRef] [PubMed] Cairns, Andrew J. G., David Blake, and Kevin Dowd. 2006. A two-factor model for stochastic mortality with parameter uncertainty: Theory and calibration. Journal of Risk and Insurance 73: 687–718. [CrossRef] Cairns, Andrew J. G., David Blake, Kevin Dowd, Guy D. Coughlan, David Epstein, Alen Ong, and Igor Balevich. 2009. A quantitative comparison of stochastic mortality models using data from England and Wales and the United States. North American Actuarial Journal 13: 1–35. [CrossRef] Cheung, Siu Lan Karen, Jean-Marie Robine, Edward Jow-Ching Tu, and Graziella Caselli. 2005. Three dimensions of the survival curve: Horizontalization, verticalization, and longevity extension. Demography 42: 243–58. [CrossRef] De Jong, Piet, and Claymore Marshall. 2007. Mortality projection based on the Wang transform. ASTIN Bulletin 37: 149–61. [CrossRef] French, Declan, and Colin O’Hare. 2014. Forecasting death rates using exogenous determinants. Journal of Forecasting 33: 640–50. [CrossRef] Fries, James F. 1980. Aging, natural death, and the compression of morbidity. New England Journal of Medicine 303: 130–35. [CrossRef] Hanewald, Katja. 2011. Explaining mortality dynamics. North American Actuarial Journal 15: 290–314. [CrossRef] Hatzopoulos, Petros, and Steven Haberman. 2015. Modeling trends in cohort survival probabilities. Insurance: Mathematics and Economics 64: 162–79. [CrossRef] HMD (Human Mortality Database). 2020. Berkeley and Rostock: University of California, Berkeley, and Max Planck Institute for Demographic Research. Available online: http://www.mortality.org/ (accessed on 28 December 2020). Hyndman, Rob J., and Anne B. Koehler. 2006. Another look at measures of forecast accuracy. International Journal of Forecasting 22: 679–88. [CrossRef] Lee, Ronald D., and Lawrence R. Carter. 1992. Modeling and forecasting U.S. mortality. Journal of the American Statistical Association 87: 659–71. [CrossRef] Lee, Ronald D., Lawrence Carter, and Shripad Tuljapurkar. 1995. Disaggregation in population forecasting: Do we need it? And how to do it simply. Mathematical Population Studies 5: 217–34. [CrossRef] [PubMed] Liu, Jia, and Jackie Li. 2019. Beyond the highest life expectancy—Construction of proxy upper and lower life expectancy bounds. Journal of Population Research 36: 159–81. [CrossRef] Medford, Anthony, and James W. Vaupel. 2019. An introduction to gevistic regression mortality models. Scandinavian Actuarial Journal 2019: 604–20. [CrossRef]
Risks 2021,9, 11 18 of 18 Nikolaevich, Khubaev Georgy. 2019. How to Increase Duration of the Population of the Country. Beijing: Scientific Research of the SCO Countries: Synergy and Integration. Niu, Geng, and Bertrand Melenberg. 2014. Trends in mortality decrease and economic growth. Demography 51: 1755–73. [CrossRef] Oeppen, Jim, and James W. Vaupel. 2002. Broken limits to life expectancy. Science 296: 1029–31. [CrossRef] Ruhm, Christopher J. 2000. Are recessions good for your Health? Quarterly Journal of Economics 115: 617–50. [CrossRef] Seklecka, Malgorzata, Norazliani Md. Lazam, Athanasios A. Pantelous, and Colin O’Hare. 2019. Mortality effects of economic fluctuations in selected eurozone countries. Journal of Forecasting 38: 39–62. [CrossRef] Tan, Chong It, Jackie Li, Johnny Siu-Hang Li, and Uditha Balasooriya. 2016. Stochastic modelling of the hybrid survival curve. Journal of Population Research 33: 307–31. [CrossRef] Wang, Shaun S. 2000. A class of distortion operations for pricing financial and insurance risks. Journal of Risk and Insurance 67: 15–36. [CrossRef] WDI (World Development Indicators). 2020. World Bank Open DataWorld Development Indicators Database. Washington, DC: The World Bank. Available online: https://data.worldbank.org/ (accessed on 28 December 2020). Wong, Chi Heem, and Albert K. Tsui. 2015. Forecasting life expectancy: Evidence from a new survival function. Insurance: Mathematics and Economics 65: 208–26. [CrossRef]