A longitudinal snalysis of the impact of distance driven on the probability of car accidents
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Boucher, Jean-Philippe; Turcotte, Roxane Article A longitudinal snalysis of the impact of distance driven on the probability of car accidents Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Boucher, Jean-Philippe; Turcotte, Roxane (2020) : A longitudinal snalysis of the impact of distance driven on the probability of car accidents, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 8, Iss. 3, pp. 1-19, https://doi.org/10.3390/risks8030091 This Version is available at: https://hdl.handle.net/10419/258044 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
risks Article A Longitudinal Analysis of the Impact of Distance Driven on the Probability of Car Accidents Jean-Philippe Boucher * and Roxane Turcotte Département de Mathématiques, Université du Québec à Montréal (UQAM), Montréal, QC H3C 3P8, Canada; [email protected] *Correspondence: boucher[email protected] Received: 20 July 2020; Accepted: 26 August 2020; Published: 1 September 2020 Abstract: Using telematics data, we study the relationship between claim frequency and distance driven through different models by observing smooth functions. We used Generalized Additive Models (GAM) for a Poisson distribution, and Generalized Additive Models for Location, Scale, and Shape (GAMLSS) that we generalize for panel count data. To correctly observe the relationship between distance driven and claim frequency, we show that a Poisson distribution with fixed effects should be used because it removes residual heterogeneity that was incorrectly captured by previous models based on GAM and GAMLSS theory. We show that an approximately linear relationship between distance driven and claim frequency can be derived. We argue that this approach can be used to compute the premium surcharge for additional kilometers the insured wants to drive, or as the basis to construct Pay-as-you-drive (PAYD) insurance for self-service vehicles. All models are illustrated using data from a major Canadian insurance company. Keywords: telematics; generalized additive models; generalized additive models for location; scale and shape; panel count data; random effects; fixed effects; distance driven 1. Introduction In the past decade, new technologies such as GPS-collected data have emerged, which offer new ways to approach car insurance pricing. Processing these data provides reliable information about drivers’ behavior. Before GPS and telematics devices, the insurance industry had to rely on proxy variables such as territory, gender and age of the drivers to measure risk. However, such covariates only describe the general behavior of insured in those groups. For example, Ayuso et al. (2016b) shows that the differences observed in claims frequency between men and women are largely attributable to vehicle use; Verbelen et al. (2018) reached a similar conclusion. In a social-political context where the use of gender in ratemaking is restricted or criticize, calculating premiums on more objective information is of interest. One piece of GPS-collected information that is directly related to the risk insured is distance driven. The relevance of including this variable in ratemaking has been studied by Ayuso et al. (2014), Ayuso et al. (2016a), Boucher et al. (2013) and Lemaire et al. (2016) among others. Boucher et al. (2017) studied the effect of distance driven and policy duration time on claim frequency and challenged the usual ratemaking practice of using contract duration as the risk exposure measure. Mileage-based pricing can generate several benefits, notably on the environment, because it encourages policyholders to reduce their annual mileage. Establishing premiums on the basis of variables that the insured can control has the significant advantage of encouraging a positive change of habit in policyholders (see for example Bolderdijk et al. (2011) and Tselentis et al. (2016)). One can argue that distance driven is correlated with other driving habits resulting from driving experience, (Ferreira and Minikel (2010)). Hence, if the model does not take this correlation into account, the resulting relationship between Risks 2020,8, 91; doi:10.3390/risks8030091 www.mdpi.com/journal/risks
Risks 2020,8, 91 2 of 19 claim frequency and the distance driven would not give an appropriate representation of how the claim frequency could change when insureds change their driving habits. This is precisely what is tackled in this paper: we indeed focus on the “marginal” effect of distance driven. The objective of our paper is not to compute a premium, but mainly to understand how the distance impacts the claim frequency when all individual characteristics of policyholders have been considered. We focus our analysis on the distance driven, yet other telematics variables could be of interest. In the study by Verbelen et al. (2018), driving time (daytime vs. nighttime) is studied along with the type of roads, while Ma et al. (2018) find that speed and acceleration affect the expected claim frequency. Ayuso et al. (2014) analyze the effect of various covariates for the time before the first crash, and compare novice and experienced drivers. More recently, Ayuso et al. (2019) propose to improve the traditional ratemaking methods by including information related to risk exposure and driving behavior of insured. Denuit et al. (2019) use predictive rating with past telematics information in a credibility model. Weidner et al. (2016) study driving behavior and vehicle use on different scales of analysis (maneuver, trip or insurance period) by means of form recognition and Fourier analysis methods. Wüthrich (2017) proposes to use speed and acceleration heat-maps to classify drivers into groups using K-means clustering. Each group is associated within a driving style and included as a categorical variable in a regression analysis. Gao and Wüthrich (2018) performed principal component analysis using singular value decomposition and bottleneck neural networks. The authors argue that a representation in two dimensions is sufficient to preserve most of the driving information, meaning that it is possible to obtain continuous representations with small-dimensional data. This representation could then be included in a Generalized Additive Model (GAM), as in the study by Gao et al. (2019). Verbelen et al. (2018) evaluate the predictive power and interpretability of telematics variables on claim frequency by comparing various types of models that include or exclude those telematics variables. The authors find that the best ratemaking structure includes both telematics and traditional covariates, while considering duration and mileage as exposure measures. In Section 2, we present the dataset used for the numerical applications throughout this work and we compare different exposure measures. In Section 3, we used a GAM Poisson, as did Boucher et al. (2017) , to link the distance driven with the number of claims. We observe the same relationship between distance and claims frequency; however we reject the “learning effect” explanation proposed by previous authors to explain the relationship, which we posit can be explained by the residual heterogeneity incorrectly captured by the underlying GAM model. Section 4presents panel count data models that are better suited to explain individual heterogeneity. In Section 5, using Generalized Additive Models for Location, Scale and Shape (GAMLSS, see Rigby and Stasinopoulos (2005) ) theory that generalizes GAM, a multivariate count distribution for all the contracts of the same insured is developed, and a penalized log-likelihood is used to estimate the parameters. In Section 6, we use another approach based on a Poisson distribution with fixed effects to account for all individual characteristics, and show that an approximately linear relationship between the distance driven and claim frequency can be found. Section 7concludes. 2. Summary of the Database The dataset that has been used for our numerical analysis comes from an important Canadian P&C insurance company. We focus our analysis on personal car insurance from the province of Ontario. In analyzing telematics data, we must be careful before jumping to general conclusions about driving behavior of the whole portfolio. Indeed, policyholders who decided to place a telematics device on their car, or to download an application on their phone that tracks all their car trips, do not correspond to the general driver population. In our case, approximately 10% to 15% of the insurance company’s portfolio chose to use the telematics option for their car insurance. Typically, these insureds correspond to one of the two following profiles:
Risks 2020,8, 91 3 of 19 1. Policyholders who are technophiles: they love new telematics technology, and want detailed information about their driving habits. Summary driving data is indeed continuously available to policyholders via a website. 2. Young and/or bad drivers. To motive policyholders to buy the telematics option, insurance companies often offer an initial discount, and the renewal discounts range from 0% to 25% depending on driving experience. 1 Because auto insurance in Ontario is very expensive and often unaffordable for some drivers, all discounts are welcome for policyholders with high insurance premiums. As a result, an unusually high proportion of risky insureds uses telematics devices or telematics app. In the dataset used, we observed the insureds for up to six insurance periods, with an average of 1.77 contracts per policyholder (see Table 1for details). Only policyholders that have been observed at least 100 days were retained for the analysis. Since this is real data, it may contain some minor irregularities. The same table shows statistics for the number of claims, where we only kept claims related to road accidents. Indeed, we wanted to study accidents related to car usage and not, for example, those caused by floods, hail, theft or vandalism. The table shows statistics for a single insured period. We note that most policyholders do not claim, that the average claim frequency for the portfolio is 6.0%, and that the maximum number of claims observed is 3. Risk Exposure Measures Table 2summarizes the statistics of various risk exposure definitions: 1. Exposure time (the time between the start and the end of the insurance contract) 2. Distance driven 3. Number of trips 4. Hours driven. Another candidate for risk exposure might be the self-reported approximation of the distance driven by the insured. However, as shown by many authors, such as Lemaire et al. (2016), the self-reported distance driven is not reliable and is often very different from the exact distance driven. Exposure time, traditionally used by insurers, would be an appropriate measure of risk exposure if every driver had about the same car usage, which is not the case. Indeed, Table 2shows that for an insured period, insureds drove between 7.1 and 76,272 km, with an average of 10,398 km. More specifically, the database also informed us about various types of car use by the insureds: 1. The maximum number of trips observed is 3317 while another one only used his car 15 times for a single insured period. 2. A policyholder drove the car for only for one hour for the whole insured period, while another driver used the car for more than 3000 h. Consequently, there are important differences between driving uses and driving habits, which justifies consideration of other measures than exposure time in the modeling. Table 1. Distribution of the number of insurance periods for the database. Number of Insurance Periods 1 2 3 4 5 6 Number of policyholders 12,562 9746 3420 844 415 11 Proportion (%) 46.5 36.1 12.7 3.1 1.5 0.0 1 Please note that it is not legally possible for an Ontario insurance company to increase the insurance premium based on the telematics information collected.
Risks 2020,8, 91 4 of 19 Table 2. Descriptive statistics for a single insured period. Average Variance Min. Max. 25th pct 50th pct 75th pct Exp. Time (in years) 0.645 0.060 0.277 1.079 0.463 0.540 0.912 Dist. Driven (in km) 10,398 55,138,376 7.1 76,272 5026 8561 13,836 Nb. of Trips 1083 383,165 15 3317 621 946 1434 Time Driven (in hours) 380 34,740 1 2159 248 356 483 Nb. of claims 0.060 0.061 0.000 3 0 0 0 Figures 1–4shows histograms of different risk exposure measures under study. Except for exposure time, every other risk exposure distribution is right-skewed. Table 2foreshadowed this result as the average was greater than the median for those risk exposures. This is another indication that some insureds make full use of their insurance time by making greater use of their car. Figures 5–8illustrate the links between claim frequency and risk exposure. A fairly clear linear trend seems to be emerging for the three non-traditional exposure measures for the first part of their respective curve, which contains most of the observations. However, we observe a strange relationship between the claims frequency and the risk exposures for higher quantiles of the distributions. We specify that each point on these graphs does not represents the same number of policyholders. Darker dots represent a larger number of policyholders. Between the three usage-based exposure measures, our choice for a more detailed analysis is distance driven. First, it seems to be the objective measure of risk of the three. Indeed, the definition of a “trip” is not clear. For example, if the driver makes a quick stop to buy gas, does it count for one or two trips because the engine stopped? For hours driven, does the time spent stopped at red lights and stuck in traffic count similarly to when the vehicle is moving? Second, it would be hard to measure exposure only according to the number of trips from a marketing point of view because those who use their vehicle only to drive short distances would probably find it unfair. 0 1000 2000 3000 4000 5000 0.00 0.25 0.50 0.75 1.00 Duration (years) Number of insured drivers Figure 1. Histogram of risk exposure (in years) Each band has a length of 0.02 year.
Risks 2020,8, 91 5 of 19 0 500 1000 1500 0 20000 40000 60000 Kilometers driven Number of insured drivers Figure 2. Histogram of distance driven (in km) Each band has a length of 500 km. 0 1000 2000 3000 4000 0 1000 2000 3000 4000 Number of trips Number of insured drivers Figure 3. Histogram of the number of trips Each band has a length of 100 trips. 0 500 1000 1500 2000 0 500 1000 1500 2000 Hours driven Number of insured drivers Figure 4. Histogram of hours driven Each band has a length of 2000 h.
Risks 2020,8, 91 6 of 19 ● ● ● ● ● ● ● ● ● ● ● ●● ● ● ● ● ● ● ● ● ●● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● 0.03 0.05 0.07 0.09 0.4 0.6 0.8 1.0 Duration (years) Claim frequency 1000 2000 3000 4000 5000 Number of insured drivers Figure 5. Claims Frequency vs. Exposure Time. ●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●● ●●●●● ●●●●●● ●● ● ● ● ●●●●●● ●● ●●●● ●● ● ● ● ●● ● ● ● ● ●● ● ● ●● ●● ● ● ● ● ● ● ●●● ● ● ● ●● ● ● ●●● ● ●● ● ● ●●● ●● ● ● ● ● ●● ● ● ● 0.00 0.25 0.50 0.75 1.00 0 20000 40000 60000 Kilometers driven Claim frequency 500 1000 1500 Number of insured drivers Figure 6. Claims Frequency vs. Distance Driven. ● ● ● ●● ●●● ●●●● ●● ● ● ● ●●●●● ● ● ● ● ● ● ● ● ● ● ● ● 0.00 0.05 0.10 0.15 0 1000 2000 3000 Number of trips Claim frequency 1000 2000 3000 Number of insured drivers Figure 7. Claims Frequency vs. Number of trips.
Risks 2020,8, 91 7 of 19 ●●●●●●●●●●●●●●●●●●●● ●●●●●●● ●●● ● ●● ● ●●●● ● ● ● ●● ● ● ● ● ●● ● ● ● ● ● ●●● ● ● ● ●● ● ● ● ● ●● ● ●●●●●●● ● ●● ● ● ● ● 0.0 0.2 0.4 0 500 1000 1500 2000 Hours driven Claim frequency 500 1000 1500 2000 Number of insured drivers Figure 8. Claims Frequency vs. Hours Driven. 3. Preliminary Risk Exposure Analysis Traditionally, the starting point for the modeling of the number of claims Ni from a policyholder i exposed to the risk of a term t is modeled by a Poisson distribution of average tλi , where λi includes classic covariates used in pricing. For a Poisson regression, t is often referred to as an offset variable. Similarly, for an insured driving a car over a distance of d km, we are seeking a model with this form of proportionality between distance driven and the expected claim frequency. Also, it could be interesting to combine t and d in the same model. To do this, one avenue is to use generalized additive models (GAM). GAMs, introduced by Hastie and Tibshirani (1986), are an extension of the generalized linear models (GLM) theory. Consequently, as for the GLM, only distributions belonging to the linear exponential family could be used as the distribution of the response variable of a GAM. In a GLM, the linear predictor for an individual i is given by g(µi) = X0 iβ , where X0 i= [x1,i , x2,i , x3,i , ... ] is a vector of covariates and β is a coefficient vector. For a GLM, the mean is given by a linear expression through a link function: GAMs relax the hypothesis of linearity, and smoothing functions s of the covariates could be included in the predictor. For example, the mean for an individual i could be given by g(µi) = s0+s1(x1,i) + s2(x2,i) + s3(x3,i) , where s0 is an intercept, sk are smoothing functions and xk,iare covariates for k∈ {1, 2, 3}. Boucher et al. (2017), by using a GAM Poisson model, analyzed the influence of duration and distance driven on the number of claims with independent cubic splines and splines with a tensor product to introduce a dependence between those two risk exposure measures (see Green and Silverman (1993) for additional details on these smoothing functions). The model with independent cubic splines is the starting point of our analysis, and we evaluate the performance of this model on our data. The model log(µi) = β0+s1(kmi) + s2(di) yields similar results to those obtained by Boucher et al. (2017) , as it can be seen in Figure 9. Indeed, we observe a strongly increasing function for the first kilometers, then it stabilizes around 40,000 km. For the higher quantile of the distribution, there are very few observations, and the confidence interval is too wide to draw conclusions. As for s2(di) , we observe a positive effect for the duration time, but no linear relationship because the function tends to stabilize. For the sake of completeness, the model with a tensor product has been fitted on our data. The tensor product includes dependency between the two exposure measures, and the fitted surface had a shape similar to that of Boucher et al. (2017). Considering the important differences between the European dataset used in Boucher et al. (2017) and our North American data, the similarity of the results is fairly interesting. First, climatic conditions are not the same given the significant
Risks 2020,8, 91 8 of 19 accumulations of snow on the ground during Canadian winters. Second, the profile of policyholders between the two databases is not the same. The Spanish data focused exclusively on young drivers while Canadian data’s profiles are more diverse as explained in Section 2. It can also be noted that the data are collected over different years for the two databases and the regulations differ from one country to another. We investigate the smoothing functions on the scale of the response level ( exp(s1(km)) and exp(s2(d)) ) to expose the multiplicative effect on λi,t , as illustrated in Figure 9. Specifically, with a log link function, we have µi,t=exp(Xi,tβ+s1(km) + s2(d)) =exp(s1(km)) exp(s2(d)) exp(Xi,tβ) =exp(s1(km)) exp(s2(d))λi,t, (1) In the study by Boucher et al. (2017), a “learning effect” is advanced to justify the look of ˆ s1(km) (and exp(ˆ s1(km)) ), where the expected number of claims seems to decrease as kilometers driven increases. We think that this effect cannot be used as an explanation. Indeed, most drivers in the insurance portfolio already have many years of driving experience. We do not think that the extra 10,000–20,000 km adds enough experience to observe a learning effect. Instead, we think that the shape of the smoothing function comes from the driver profiles: the lower quantiles of the distribution of the distance driven does not come from the same (type of) drivers as the higher quantiles. This means that models based on Figure 9cannot be used to understand the relationship between the distance driven and the number of claims, and might not be used to set the premium for insured that suddenly change their driving habits, because it does not nearly tell us how their risk is changing. As an example to illustrate the situation, we can suppose an insured who suddenly decides to drive 50,000 km instead of 40,000 km. Based on Figure 9, we would expect a decrease in the expected claims frequency. This is however impossible: the number of claims in the first 40,000 km cannot change, and the extra 10,000 km can only add other claims. In other words, if insureds choose to drive their cars rather than leaving it at home, the risk should always be greater. The slope could change as distance increases, but it should always be strictly positive since the risk is greater, meaning that the smoothing function (as the one observed in Figure 9) should always be increasing. Our results, and those of Boucher et al. (2017), do not show a strictly positive relationship between claim number and distance driven. We think that this can be explained by the residual individual heterogeneity of the model, which the basic Poisson GAM does not seem to capture correctly. One explanation comes from the fact that GAM supposes independence between all contracts of the same insured. We think that a more general model that relaxes this assumption should be used to correctly measure the impact of the distance driven on the risk of accidents. 1 2 3 0 20000 40000 60000 80000 Kilometers driven (km) exp(s(km)) 0.8 1.0 1.2 0.4 0.6 0.8 1.0 Duration (year) exp(s(year)) Figure 9. exp(ˆ s1(km)) and exp(ˆ s2(year)) from the Poisson GAM estimated with Canadian data.
Risks 2020,8, 91 15 of 19 beyond this point, which is not very significant. What has been called the “learning effect”, observed in Section 3, has disappeared and we observe a much more logical and coherent relationship between distance traveled and frequency than before. The relationship between claim frequency and the distance driven should be understood as the marginal impact of each additional kilometer driven or not-driven. Explicitly, as we approximated exp(s(km)) by 0.25 +1 15,000 kmi,t (the red line in Figure 12), we then have Nit ∼Poisson (exp(αi)exp(s(km))) ∼Poisson (exp(αi)(a+b kmi,t)) ∼Poisson 0.25 exp(αi) + 1 15,000 exp(αi)kmi,t. We see that the slope, i.e., the marginal impact of each additional kilometer driven or not-driven, is not the same for each insured because it depends on αi . To illustrate this difference, we use the estimated values of αi for several insureds. Figure 13 shows the relationship between claim frequency and distance driven for different individuals (the policyholder with the minimum, maximum, median, 25th and 75th percentile individual parameter value). With this model, we then reconcile the intuition that each kilometer should increase the risk for an individual, but that this increase could be different for each driver. In summary, instead of referring to the “learning effect” to understand the left-hand graph of Figure 9, we should understand instead that typical insureds who drive more than 60,000 km per year are better risks per kilometer than insureds who drive approximately 40,000 km per year. That obviously does not mean that insureds that drive 40,000 km per year should drive 60,000 km to reduce their risk. The difference between insureds related to their risk per kilometer can be explained by many factors: more frequent use of the highway, higher proportion of driving outside rush hours, etc. However, for each driver, independently of their driving risk per kilometer, the risk of an accident will always increase for each additional kilometer driven (by approximately 1 15,000 ). To conclude about the fixed effects model, note that the risk is still present even when the driving distance is zero. This is counter-intuitive because we can presume that someone who does not drive at all should have an expected claim frequency of zero. We agree. However, the real risk exposure is never completely null and the intercept could represent situations where an accident is possible even without driving a lot (e.g., it may occur very close to the insured’s home). Moreover, even if the car would never actually be used, hit and run situations are also possible. 0 5 10 15 0 20000 40000 60000 Distance driven (km) Claim number −2 −1 0 1 Individual parameter Figure 13. Exposure measure for different individual parameters.
Risks 2020,8, 91 16 of 19 6.4. Which Effect Should Be Used in Practice? Random and fixed effects seem to generate contradictory results, and we may wonder which model we should then use in practice, particularly for ratemaking. This has already been discussed in the actuarial literature by Boucher and Denuit (2006), but it is worth reexamining it in the context of telematics data, for distance driven in our case. First, the fixed effects model is more general than the random effects model, which means that in case of contradictory results, fixed effects should always be preferred. Equation (3) can be derived as: Pr[Ni,1 =ni,1, ..., Ni,T=ni,T] =ˆ∞ 0 Pr[Ni,1 =ni,1, ..., Ni,T=ni,T|xi,1, ..., xi,T,αRE i]f(αRE i|xi,1, ..., xi,T)dαRE i =ˆ∞ 0 T ∏ t=1 Pr[Ni,t=ni,t|xi,1, ..., xi,T,αRE i]!f(αRE i)dαRE i =ˆ∞ 0 T ∏ t=1 exp(−αRE iλRE i,t)(αRE iλRE i,t)ni,t ni,t!!f(αRE i)dαRE i We can see that we have to suppose an additional assumption: from the first to the second line of development, f(αRE i|xi,1 , ..., xi,T) becomes f(αRE i) . That means that we must suppose that random effects are independent of observed covariates. Empirical analyses have shown that this is not the case. Indeed, as shown by Boucher and Denuit (2006), random effects do not have the same distribution for young drivers as for older ones, and depends on gender, for example. However, this is a typical assumption made in actuarial science, and Boucher and Denuit (2006) discusses the consequences of not satisfying this assumption. The authors concluded that the interpretation of random effects results are tricky. On the other hand, fixed effects modeling, even if theoretically better, is not amenable to ratemaking: • The model requires evaluating an individual parameter αi for each insured i in the portfolio. This raises a problem for new policyholders. •For a small value of Ti,bαimay be incorrectly estimated. • As the model estimates each individual αi as ni,• λi,• , policyholders without claims will have an expected number of claims of 0, meaning that the premium of these insureds should be zero. As Boucher and Denuit (2006) conclude for basic ratemaking purposes, even if theoretically problematic, the random effects model should be preferred over a fixed effects model: random effects are flexible enough to compute premiums for new insureds, and do not generate a premium of 0 for insureds without claims. However, actuaries must understand that the parameters obtained by random effects models only indicate the apparent effect of the covariates, and not a causal effect (or what might be call the real impact). To compare the fixed effects results with those of the random effects model, the approximate relationship for the median value of the individual parameter bαi has been plotted over the smoothed function of distance traveled of the random effect approach (see Figure 14). Interestingly, the two curves are similar. Regarding the use of the results of a fixed effects model, fixed effects should be used to understand the “true” relationship between covariates and claims experience. For ratemaking, fixed effects should be used to compute the premium surcharge for each additional kilometer the insureds drive. In our case, it represents an increase of bαi1 15,000 per km, for claim frequency. Using this approach, insurers will avoid the situation where an insured could see a premium reduction if, for example, he decides to drive 50,000 km instead of 40,000 km, as we saw with a basic GAM approach. Fixed effects can be used
Risks 2020,8, 91 17 of 19 to construct PAYD insurance solely based on kilometers driven for self-service vehicles, where drivers’ profile cannot be directly used for ratemaking. Research is required in this area. 0 1 2 3 0 20000 40000 60000 80000 Kilometers driven (km) exp(s(km)) Figure 14. Comparison between the random effect approach and the fixed-effect approach for the median value of the individual parameter. 7. Conclusions We have studied the relationship between claim frequency and the distance driven through different models by observing smooth functions. We first reproduced with our data the model proposed by Boucher et al. (2017) and observed what the authors called the “learning effect,” where the expected number of claims seems to decrease as kilometers driven increase. Given that most drivers in the insurance portfolio already has many years of driving experience, we rejected the conclusion that an additional 10,000–20,000 km adds enough experience to observe a learning effect. Instead, we supposed that the residual heterogeneity was incorrectly captured by the underlying GAM model. We then evaluated panel data models with fixed and random effects. Using GAMLSS theory, which generalizes GAM, a multivariate count distribution for all the contracts of the same insured was developed. Smoothing functions were added in the mean parameter of the multivariate distribution, and a penalized log-likelihood was used to estimate the parameters. A grid of penalties, generating more than 1000 MVNB, was used to find the best distribution. However, again, the fitted smoothed function for the distance driven by the Poisson distribution with random effects did not seem to correctly describe the relationship between distance and claim frequency. Indeed, the expected number of claims still decreases disproportionately with kilometers driven. We then used the Poisson with fixed effects to account for all individual characteristics. Because Poisson with fixed effects can be estimated by using covariates that identify each insured, we show that a simple GAM model without intercept can be used to include a smoothed function in the mean parameter. We then observed an approximately linear relationship between the distance driven and claim frequency when all individual characteristics have been accounted for in an individual parameter. This unravels the potential for the distance traveled as an exposure variable, even though this variable could not serve as a rating model. However, we think that the model proposed can be used to compute the premium surcharge for additional kilometers the insured wants to drive, or as the basis to construct PAYD insurance for self-service vehicle. The new telematics data available in automobile insurance offers several new challenges. These data increase the possibility of identifying factors that make accidents more probable. Models like
Risks 2020,8, 91 18 of 19 the fixed effects models proposed in the paper, make it possible to better capture the real effect of a covariate on risk. By using various models that do more than predict or calculate the insurance premium, research by insurers could shed light on risk in auto insurance. We therefore believe that many statistics compiled by telematics devices could be studied from such an angle in the future. Author Contributions: The authors worked together on all aspects of the work. All authors have read and agreed to the published version of the manuscript. Funding: Jean-Philippe Boucher and Roxane Turcotte gratefully acknowledge the financial support of Cooperators General Insurance Company through the Co-operators Chair in Actuarial Risk Analysis. The authors would also like to thank the financial support from the Natural Sciences and Engineering Research Council of Canada. Acknowledgments: Jean-Philippe Boucher and Roxane Turcotte thank Gabriel Alepin for his help in the code of the GAMLSS model. Conflicts of Interest: The authors declare no conflict of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript, or in the decision to publish the results. References Ayuso, Mercedes, Montserrat Guillén, and Ana María Pérez-Marín. 2014. Time and distance to first accident and driving patterns of young drivers with pay-as-you-drive insurance. Accident Analysis & Prevention 73: 125–31. Ayuso, Mercedes, Montserrat Guillen, and Ana María Pérez Marín. 2016a. Using gps data to analyse the distance travelled to the first accident at fault in pay-as-you-drive insurance. Transportation Research Part C: Emerging Technologies 68: 160–67. [CrossRef] Ayuso, Mercedes, Montserrat Guillen, and Ana María Pérez-Marín. 2016b. Telematics and gender discrimination: Some usage-based evidence on whether men’s risk of accidents differs from women’s. Risks 4: 10. [CrossRef] Ayuso, Mercedes, Montserrat Guillen, and Jens Perch Nielsen. 2019. Improving automobile insurance ratemaking using telematics: Incorporating mileage and driver behaviour data. Transportation 46: 735–52. [CrossRef] Bolderdijk, Jan Willem, Jasper Knockaert, E. M. Steg, and Erik T. Verhoef. 2011. Effects of pay-as-you-drive vehicle insurance on young drivers’ speed choice: Results of a dutch field experiment. Accident Analysis & Prevention 43: 1181–86. Boucher, Jean-Philippe, and Michel Denuit. 2006. Fixed versus random effects in poisson regression models for claim counts: A case study with motor insurance. ASTIN Bulletin: The Journal of the IAA 36: 285–301. [CrossRef] Boucher, Jean-Philippe, Ana Maria Pérez-Marín, and Miguel Santolino. 2013. Pay-as-you-drive insurance: The effect of the kilometers on the risk of accident. In Anales del Instituto de Actuarios Españoles. 19 vols. Madrid: Instituto de Actuarios Españoles, pp. 135–54. Boucher, Jean-Philippe, Steven Côté, and Montserrat Guillen. 2017. Exposure as duration and distance in telematics motor insurance using generalized additive models. Risks 5: 54. [CrossRef] Cameron, A. Colin, and Pravin K. Trivedi. 2013. Regression Analysis of Count Data. 53 vols. Cambridge: Cambridge University Press. Denuit, Michel, Montserrat Guillen, and Julien Trufin. 2019. Multivariate credibility modelling for usage-based motor insurance pricing with behavioural data. Annals of Actuarial Science 13: 378–99. [CrossRef] Denuit, Michel, Xavier Maréchal, Sandra Pitrebois, and Jean-François Walhin. 2007. Actuarial Modelling of Claim Counts: Risk Classification, Credibility and Bonus-Malus Systems. Hoboken: John Wiley & Sons. Eilers, Paul H. C., and Brian D. Marx. 1996. Flexible smoothing with b-splines and penalties. Statistical Science 11: 89–102. [CrossRef] Ferreira, Joseph, and Eric Minikel. 2010. Pay-as-You-Drive Auto Insurance in Massachusetts: A Risk Assessment and Report on Consumer, Industry and Environmental Benefits. Boston: Conservation Law Foundation. Gao, Guangyuan, and Mario V. Wüthrich. 2018. Feature extraction from telematics car driving heatmaps. European Actuarial Journal 8: 383–406. [CrossRef] Gao, Guangyuan, Shengwang Meng, and Mario V. Wüthrich. 2019. Claims frequency modeling using telematics car driving data. Scandinavian Actuarial Journal 2019: 143–62. [CrossRef]
Risks 2020,8, 91 19 of 19 Green, Peter J., and Bernard W. Silverman. 1993. Nonparametric Regression and Generalized Linear Models: A Roughness Penalty Approach. Boca Raton: Chapman and Hall/CRC. Hastie, Trevor, and Robert Tibshirani. 1986. Generalized additive models. Statistical Science 1: 297–310. [CrossRef] Inouye, David I., Eunho Yang, Genevera I. Allen, and Pradeep Ravikumar. 2017. A review of multivariate distributions for count data derived from the poisson distribution. Wiley Interdisciplinary Reviews: Computational Statistics 9: e1398. [CrossRef] [PubMed] Lemaire, Jean, Sojung Carol Park, and Kili C. Wang. 2016. The use of annual mileage as a rating variable. ASTIN Bulletin 46: 39–69. [CrossRef] Ma, Yu-Luen, Xiaoyu Zhu, Xianbiao Hu, and Yi-Chang Chiu. 2018. The use of context-sensitive insurance telematics data in auto insurance rate making. Transportation Research Part A: Policy and Practice 113: 243–58. [CrossRef] Molenberghs, Geert, and Geert Verbeke. 2006. Models for Discrete Longitudinal Data. Berlin: Springer Science & Business Media. Rigby, Robert A., and D. Mikis Stasinopoulos. 2005. Generalized additive models for location, scale and shape. Journal of the Royal Statistical Society: Series C (Applied Statistics) 54: 507–54. [CrossRef] Tselentis, Dimitrios I., George Yannis, and Eleni I. Vlahogianni. 2016. Innovative insurance schemes: Pay as/how you drive. Transportation Research Procedia 14: 362–71. [CrossRef] Verbelen, Roel, Katrien Antonio, and Gerda Claeskens. 2018. Unravelling the predictive power of telematics data in car insurance pricing. Journal of the Royal Statistical Society: Series C (Applied Statistics) 67: 1275–304. [CrossRef] Weidner, Wiltrud, Fabian W. G. Transchel, and Robert Weidner. 2016. Classification of scale-sensitive telematic observables for riskindividual pricing. European Actuarial Journal 6: 3–24. [CrossRef] Wood, Simon N. 2017. Generalized Additive Models: An Introduction with R. Boca Raton: Chapman and Hall/CRC. Wüthrich, Mario V. 2017. Covariate selection from telematics car driving data. European Actuarial Journal 7: 89–108. [CrossRef] c 2020 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).