Profit uplift modeling for direct marketing campaigns: approaches and applications for online shops
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Baier, Daniel; Stöcker, Björn Article — Published Version Profit uplift modeling for direct marketing campaigns: approaches and applications for online shops Journal of Business Economics Provided in Cooperation with: Springer Nature Suggested Citation: Baier, Daniel; Stöcker, Björn (2021) : Profit uplift modeling for direct marketing campaigns: approaches and applications for online shops, Journal of Business Economics, ISSN 1861-8928, Springer, Berlin, Heidelberg, Vol. 92, Iss. 4, pp. 645-673, https://doi.org/10.1007/s11573-021-01068-3 This Version is available at: https://hdl.handle.net/10419/287419 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) Journal of Business Economics (2022) 92:645–673 https://doi.org/10.1007/s11573-021-01068-3 1 3 ORIGINAL PAPER Profit uplift modeling fordirect marketing campaigns: approaches andapplications foronline shops DanielBaier1 · BjörnStöcker2 Accepted: 20 October 2021 / Published online: 19 November 2021 © The Author(s) 2021 Abstract In order to select “best” customers for a direct marketing campaign, response models are widespread: a sample of customers receives an ad, a catalog, a sample pack, or a discount offer on a test basis. Then, their responses (e.g., website visits, conversions, or revenues) are used to build a predictive model. Finally, this model is applied to all customers in order to select “best” ones for the campaign. However, up to now, only models that reflect website visits, conversions, or revenues have been proposed. In this paper, we discuss the shortcomings of these traditional approaches and propose profit uplift modeling appoaches based on one-stage ordinary regression and random forests as well as two-stage Heckman sample selection and zero-inflated negative binomial regression for parameter estimation. The new approaches demonstrate superiority to the traditional ones when applied to real-world datasets. One dataset reflects recent discount offers of a large online fashion retailer. The other is the wellknown Hillstrom dataset that describes two Email campaigns. Keywords Uplift modeling· Heckman sample selection model· Zero-inflated negative binomial regression· Random forests· Online shops JEL Classification C01· C53· M31· M37 * Daniel Baier [email protected] Björn Stöcker [email protected] 1 Chair ofMarketing andInnovation, University ofBayreuth, Universitätsstraße 30, 95447Bayreuth, Germany 2 Head ofCRM, BAUR Versand, Bahnhofstraße 10, 96224Burgkunstadt, Germany
646 D.Baier, B.Stöcker 1 3 1 Introduction Before launching a direct marketing campaign, often, a sample of customers is testwise contacted. Their desired responses (e.g., website visits, conversions, revenues) as well as their past information and buying behavior is used to build a response model. Then, this model is applied to all customers and to select likely responders for the campaign. However, this traditional approach has two shortcomings: • First, response models focus on likely responders, possibly independent of the contact. This could be a waste of money, e.g. in case of unnecessarily distributed sample packs or discount offers. • Second, up to now, response models only predict binary outcomes (website visits, conversions) or revenue outcomes, not the more informative profit outcomes at the individual level. Both shortcomings restrict the usefulness of the traditional approaches for maximizing profit. In this paper, we propose new profit uplift modeling approaches as alternatives: first, uplifts focus—in contrast to responses—on the incremental response due to a treatment using control groups of customers. Second, profit is more difficult to model since this outcome is only observable in a few cases but more closely related to the main objective than website visit, purchase, or revenue. The proposed new approaches in this paper extend findings from the field of binary and revenue uplift modeling (e.g., Radcliffe and Surry 1999, 2011; Kane etal. 2014; Rudaś and Jaroszewicz 2018; Gubela etal. 2020) and from the field of two-stage estimation via sample selection (see, e.g., Heckman 1979) and zero-inflated regression (see, e.g., Lambert 1992; Ridout et al. 2001) as well as one-stage parameter estimation via ordinary regression and random forest. We show that the new approaches are well suited to select “best” customers as targets for direct marketing campaigns and improve profit. The paper is organized as follows: In Sect. 2, we discuss the traditional approaches and their shortcomings and, in Sect.3, the new approaches. In Sect.4, the superiority of the new over the traditional approaches is demonstrated using a new dataset from a major German online retailer (with a sample of n = 155,388 customers). In Sect.5, the well-known Hillstrom direct marketing campaign dataset (with a sample of n = 64,000 customers) is used for the same purpose. The paper closes with conclusions and outlook. 2 Background andrelated work Testing and predictive modeling are assumed to be the analytical cornerstones of today’s direct marketing (Blattberg etal. 2008). The modeling process usually consists of the following five steps: (1) Define the managerial problem in terms of
647 1 3 Profit uplift modeling fordirect marketing campaigns:… a campaign and its intended effects, (2) translate this description to a predictive model with treatment, responses, and potential predictors, (3) sample customers for collecting responses, (4) calibrate and validate the predictive model, (5) apply the model to all customers and select “best” customers according to the predictive model. Typical managerial problems are selecting targets for an acquisition campaign at hand, deciding on customers to receive a catalog or inlay, or identifying promising customers for a customer tier program. Outcomes are the response to the treatment, predictors are customer characteristics (age, gender, income if available) and variables that describe past information and buying behavior in the customer database (see, e.g., Blattberg etal. 2008 for an overview). Let Yi be the binary ( Yi∈{0, 1} ) or continuous ( Yi∈ℝ ) outcome for customer i in the customer sample, 𝐱i ( 𝐱i∈ℝm ) customer i’s values for the m predictors, and 𝜏i the indicator whether customer i received the treatment (= 1) or not (= 0). Then, the main goal for a traditional response modeling approach would be to predict the following purchase likelihood (in case of binary outcomes) or scalar values (in case of continuous outcomes) as customer scores for selecting targets, The data of the treated customers ( 𝜏i=1 ) are used to calibrate the response model (using, e.g., logistic regression or simple regression depending on the scale of the outcome). Then, the whole customer database is used for prediction. The customers with the highest (response) scores are targets for the campaign. However, this response modeling approach has one major shortcoming: it favors customers who respond most likely, but it does not take into account that some of them would also respond if not treated. When the treatment is a discount, a voucher, a catalog, an inlay, or one has to deal with postage, this could result in a waste of money for the company. Therefore, recently, uplift models have been proposed: responses are again collected from a sample of treated customers (the treatment group), but also from a sample of not treated customers (the control group). An uplift model now predicts the difference in the response of a customer if treated ( 𝜏i=1 ) and if not treated ( 𝜏i=0 ), the so-called uplift score, Terms like differential response (e.g., Radcliffe and Surry 1999), true lift (Lo 2002), or uplift (Radcliffe and Surry 2011) are used for the same idea. Formula (2) allows to estimate the effect of the treatment and enables the company to select customers where the treatment has the highest impact. However, when trying to estimate the model parameters, a problem arises from the fact that per customer, only one of these two responses is observable: a customer is part of the treatment group ( 𝜏i=1 ) or part of the control group ( 𝜏i=0 ), not of both groups. Consequently, an uplift model cannot be estimated directly when using formula (2). Instead, one straightforward idea is to develop two separate models (the so-called two model approach): a first model is derived similar to formula (1), based on the treatment group. This model predicts the outcome (1) Responsei( 𝐱 i)=E(Yi| 𝐱 i). (2) Uplifti (𝐱 i )=E(Y i| 𝐱 i, 𝜏 i = 1 )−E(Y i| 𝐱 i, 𝜏 i = 0 ) .
648 D.Baier, B.Stöcker 1 3 if treated in terms of 𝐱i for all customers. A second model is derived using the control group. This model predicts the outcome if not treated in terms of 𝐱i for all customers (see Radcliffe and Surry 1999), the difference between the predictions of the two models is the uplift. An alternative solution is the so-called interaction model proposed by Lo (2002): an interaction (response) model uses the treatment ( 𝜏i ) and interactions between the predictors and the treatment as additional predictors. The interaction model can be calibrated on the treatment and the control group simultaneously. Then, for all customers, predictions for all customers are derived via formula (2) by setting the treatment for all customers to 1 in the first term and 0 in the second term. Over the years, a remarkably high number of uplift modeling approaches, including algorithms to estimate their parameters, have been proposed. Table1 gives an overview. As one can easily see, most of them aim at predicting uplifts for binary outcomes, e.g., indicators for a visit, conversion, or purchase. Here, logistic regression or decision trees can be applied to estimate model parameters. However, more recently, also revenue uplift modeling approaches have become popular (Gubela etal. 2020; Rudaś and Jaroszewicz 2018). The main idea behind this new development is that the revenue uplift more closely relates to economic objectives than a website visit or purchase uplift. However, in the next section, we will see that even revenue uplift modeling approaches sort customers suboptimally. Another interesting aspect in Table 1 is that most recent uplift modeling approaches rely on transformed outcomes for parameter estimation. This transformation was introduced for binary outcomes in a seminal paper (Lai 2006) and later extended to continuous outcomes (Gubela etal. 2020; Rudaś and Jaroszewicz 2018). The main idea behind it is to transfer as much information as possible from the observed responses in the two groups into the dependent variable and so being able to directly estimate the uplift model parameters (a so-called direct model). So, e.g., Rudaś and Jaroszewicz (2018)—following the proposal of Lai (2006) for binary outcomes—proposed to estimate their revenue uplift model, directly using transformed revenue outcomes, for parameter estimation. qT and qC are the fractions of the treatment group and the control group in the customer sample. Rudaś and Jaroszewicz (2018) discuss in their paper that this weighting facilitates unbiased estimation of the model parameters when relying on linear models. The main idea behind the positive weighting of the observed revenues in the treatment sample and the negative weighting of the observed revenues in the control sample is that so the best possible information is forwarded to parameter estimation. It is assumed that the purchasers in the treatment group generate probably a (low to high) positive revenue uplift. Likewise, it is (3) Uplifti( 𝐱 i)=E(Zi| 𝐱 i), (4) Z i= ⎧ ⎪ ⎨ ⎪ ⎩ +1 qTYiif 𝜏i=1∧Yi>0 0if Yi=0 −1 qCYiif 𝜏i=0∧Yi>0 ,
649 1 3 Profit uplift modeling fordirect marketing campaigns:… Table 1 Uplift modeling approaches Only new approaches are reflected, CART classification and regression tree, DT decision tree, MLP multilayer perceptron, RF random forest, sel. selection, SVM support vector machine Approach Outcome Algorithm Reference Differential response analysis: modeling true response Binary DT Radcliffe and Surry (1999) Incremental value modeling Binary DT Hansotia and Rukstales (2002) The true lift model Binary Logistic regression Lo (2002) Influential marketing: a new direct marketing strategy Binary (transformed) Association rules, DT, Logistic regression Lai (2006) Using control groups to target on predicted lift Continuous CART Radcliffe (2007) Uplift modelling with significance-based uplift trees Continuous CART Radcliffe and Surry (2011) DTs for uplift modeling with multiple treatments Multiple binary DT Rzepakowski and Jaroszewicz (2012) Support vector machines for uplift modeling Binary (transformed) SVM Zaniewicz and Jaroszewicz (2013) Uplift random forests Binary (transformed) Causal conditional inference tree/RF Guelman etal. (2015) Mining for truly responsive customers Binary (transformed) Logistic regression Kane etal. (2014) Ensemble methods for uplift modeling Binary (transformed) Ensemble methods Sołtys etal. (2015) Lp-support vector machines for uplift modeling Binary (transformed) SVM Zaniewicz and Jaroszewicz (2017) Revenue uplift modeling Continuous (transformed) Linear regression Rudaś and Jaroszewicz (2018) Revenue uplift modeling Continuous (transformed) Lasso, Ridge, and TheilSen regression, MLP, RF Gubela etal. (2020) Profit uplift modeling Continuous (transformed) OLS, Heckman sample sel., Zero-inflated NB, RF This paper
650 D.Baier, B.Stöcker 1 3 assumed that the purchasers in the control group generate probably a (low to high) negative revenue uplift. Another major problem with uplift modeling approaches is to validate their predictions at the customer level as for these predictions—as mentioned above—no observations exist. The widespread solution for this problem is to develop so-called Qini curves and calculate the so-called Qini coefficient Q (Radcliffe 2007; Radcliffe and Surry 2011): the customers are sorted according to a descending (uplift) score and partitioned into deciles (or other partitions of the customers) with similar scores. Then, within the deciles, customer responses from the treatment group are averaged as well as customer responses from the control group. The difference of these two means is assumed to be the “observed” uplift in this decile. Figure1 shows the typical results for such a validation applied to (uplift) scores from a sample dataset. In both diagrams, the customers are sorted according to descending uplift predictions from left to right and grouped in deciles. In the right diagram, the calculated average (“observed”) uplift per decile is given, as discussed above. In the left diagram, from decile to decile, the average cumulative uplift is plotted, which means that for the first decile, the values in the left and right diagram are identical, but from then, aggregated values up to the current decile are given in the left diagram. The last value (here: 0.045) of this so-called Qini curve reflects the uplift across all ten deciles (all customers of the treatment and the control group) which is identical for all scorings based on the data. For comparisons, also the Qini curve for a random sorting (the random uplift model) is plotted in the left diagram. Its incremental uplift curve connects the zero point with the average uplift across all customers, the last value (0.045). The quality of an uplift model is judged by its ability to sort customer deciles according to decreasing “observed” uplifts (in the right diagram) but—similar to ROC curves—also by calculating the area between the Qini curve and the line for the random model (in the left diagram). Here the length of the x-axis Fig. 1 Qini curve (left) and mean uplifts (right) for a sample dataset. The area between the Qini curve of an uplift model and a random model is the Qini coefficient Q (here: Q = 0.0296), which can be used for model selection
651 1 3 Profit uplift modeling fordirect marketing campaigns:… is assumed to be 1. In Fig.1, this value—the so-called Qini coefficient Q—is 0.0296 and could serve for comparisons with other uplift models (other sortings of customers according to their predicted scores). The random model has Q = 0, the maximum is data-dependent. It should be noted that these two diagrams can be generated for uplift models with binary response outcomes but also for uplift models with continuous outcomes (as in revenue uplift modeling or our new profit uplift modeling approach discussed in the following section). 3 Profit uplift modeling approaches foronline shops 3.1 Potential usefulness ofprofit uplift modeling approaches As already discussed, most uplift modeling approaches reflect binary outcomes. Only recently, continuous outcomes have received more interest, e.g., in the papers by Rudaś and Jaroszewicz (2018) as well as Gubela etal. (2020). This is surprising since, from the beginning of the development of uplift modeling approaches, also datasets with continuous outcomes have been made available. So, e.g., the famous Hillstrom dataset (Radcliffe 2008)—which is often seen as the standard dataset in uplift modeling and has been used in many papers when uplift models were introduced or compared—contains as binary outcomes the website visits (= 1: yes, = 0: no) and the purchase information (= 1: yes, = 0: no) but also the revenue generated by this purchase (spend in $). However, maybe since the share of purchasers in this dataset (0.9% of the customers) and the revenue uplift were very low, and, additionally, the revenues concentrate on very few purchasers, this dataset did not stimulate the scientific community to develop continuous outcome uplift models. Even in the newer and methodologically advanced paper Rudaś and Jaroszewicz (2018), this dataset is only used as a basis for a simulation at the end of the paper. The main methodological progress in revenue uplift modeling in their paper was demonstrated by using synthetic data. However, recently, Gubela etal. (2020) have demonstrated in their paper with large real-world datasets (nearly 3 million sessions from visits at 25 European online shops) that revenue uplift modeling approaches provide further insights. This superiority of a continuous outcome uplift modeling approach can also be seen when reflecting the assumed behavior of a small sample of customers as shown in Table2. Here, for 12 customers, their potential outcomes (purchases, revenues, and profits) are given in case of a direct marketing campaign with a discount offer of d = 20% and a profit margin of m = 30%. Profits are calculated for customers in the treatment group as 10% (= m–d) and for the control group as 30% (= m) of the revenue. One can easily see that the 12 customers reflect a typical behavior: they show— on average—a purchase outcome uplift (8%) when offered a discount, they generate a higher revenue when a discount is offered (+ 69 €), but it is not useful to offer the discount to all customers since the profit uplift—on average—is negative (–9 €). Only five customers (1, 2, 3, 4, and 5) show a profit uplift, which means that only these five customers should be offered the discount. The customer sorting according
652 D.Baier, B.Stöcker 1 3 to the purchase outcome and the revenue outcome differs: the three customers with the highest revenue uplift show a purchase uplift of 0. However, as also can be seen in Table2, both sortings considerably differ from the sorting according to the profit uplift: if the customers were targeted according to their revenue uplift, customers with a positive profit uplift but also with a negative profit uplift would receive a discount offer. It should be noted that this difference in sorting heavily relies on the ability of discounts to generate additional revenues but also on the fact that in online shops, high discounts are widespread but would lead to losses if granted to all customers. Moreover, it should be mentioned that Table2 reflects an ideal situation inso-far that from each customer, two observations are available—the outcomes with and without treatment—which in reality is not possible. 3.2 Profit uplift modeling approaches indetail After demonstrating the potential usefulness of profit uplift modeling approaches, now, they are discussed in detail. The main idea is to use formulae (2) (as a two model or an interaction model approach) or (3) and (4) (as a direct approach) for modeling continuous outcomes but to replace the observed revenue by derived profits and the revenue uplift predictions by profit uplift predictions. We follow Blattberg etal. (2008) as in Sect.2 and discuss the five steps of the predictive modeling process now in detail: Table 2 Sample of customers with potential purchase, revenue, profit if treated (offered a discount of 20% at a margin of 30%) and if not treated (no discount offer) Interpretation: customer 1 generates a revenue of 160 € if treated and 0 € if not treated. The treatment generates a revenue uplift of 160 € (= 160–0 €) and a profit uplift of 16 € (= 160 € * (30–20%) – 0 € * 30%). Please note that this perfect information is not observable in practice since a customer can only be part of the treatment group (if treated) or the control group (if not treated) Cus-tomer Purchase Revenue Profit If treated If not tr. Uplift If treated If not tr. Uplift If treated If not tr. Uplift 1 1 0 1 160 € 0 € 160 € 16 € 0 € 16 € 2 1 1 0 300 € 60 € 240 € 30 € 18 € 12 € 3 1 0 1 40 € 0 € 40 € 4 € 0 € 4 € 4 1 0 1 30 € 0 € 30 € 3 € 0 € 3 € 5 1 0 1 20 € 0 € 20 € 2 € 0 € 2 € 6 1 1 0 70 € 40 € 30 € 7 € 12 € –5 € 7 0 1 –1 0 € 20 € –20 € 0 € 6 € –6 € 8 0 1 –1 0 € 40 € –40 € 0 € 12 € –12 € 9 0 1 –1 0 € 60 € –60 € 0 € 18 € –18 € 10 1 1 0 400 € 200 € 200 € 40 € 60 € –20 € 11 1 1 0 500 € 250 € 250 € 50 € 75 € –25 € 12 1 1 0 270 € 290 € –20 € 27 € 87 € –60 € Mean 75% 67% 8% 149 € 80 € 69 € 15 € 24 € –9 €
659 1 3 Profit uplift modeling fordirect marketing campaigns:… Table 4 472 variables of the BAUR dataset that describe past information and buying behavior Variable category Number of variables Description Recency 23 Variables that count days since last order (w.r.t. discount types, item categories, and time slots) Frequency 193 Variables that count past orders (w.r.t. discount types, item categories, and time slots) Monetary value 191 Variables that reflect past revenues (w.r.t. discount types, item categories, and time slots) Shop visit 14 Variables that describe the online information behavior (w.r.t. number of visits, visit duration, basket size and value across time slots and item categories) Sensitivity to recommendations 3 Variables that describe the number of orders and their value due to recommendations (w.r.t. time slots) Sensitivity to discounts 26 Variables that describe the share of orders with discounts to all orders in the past (w.r.t. discount types, item categories, and time slots) Return behavior 22 Variables that describe the number of returns and their value (w.r.t. time slots)
660 D.Baier, B.Stöcker 1 3 • Heckman: the two-stage Heckman selection model (Heckman 1979) is estimated based on the binary outcome (purchase) and—in case of a predicted purchase— on the profit uplift. For parameter estimation, first, the observed profit for all purchasers is derived from the observed revenue by multiplying with the margin (m = 30%) for the purchasers in the control group and with the margin minus discount (m–d = 10%) for the purchasers in the treatment group. Then, the profit response is transformed to “observed” profit uplift according to formula (4), and the Heckman selection model is estimated. Finally, profit uplift predictions can be directly derived for all customers using formula (3) via formulae (5) and (6). Besides this direct model approach (using the “observed” profits for estimation) also a two model approach according to formula (2) was used (using the profit responses in the treatment and control group for separate estimations). For all estimations, the R package and R function sampleSelection was applied. • OLS and RF: as one-stage models, simple regression (OLS) and random forest (RF) (Breiman 2001) is used. We apply glm from the MASS package in R in case of OLS and the ranger implementation in R (Wright and Ziegler 2017) in case of RF to the “observed” profit uplifts as a direct modeling approach. Again, predictions for the profit uplift outcome can be derived for all customers directly according to formula (3). • Zeroinfl: the two-stage zero-inflated Poisson regression model (Lambert 1992) and its zero-inflated negative binomial regression model alternative (Ridout etal. 2001) assume non-negative count data as input. Therefore, first, the observed profit has to be converted to Millicent (to preserve variability) and to be rounded. Also, as discussed in Sect.3, an interaction model is useful (as an alternative to the two model approach) that includes the treatment indicator (1 for customers in the treatment group, 0 for the others) and its interactions with the other predictors. The estimated interaction model then is used for predicting the profit uplift as the difference between the predicted profit when the treatment indicator is set to 1 and the predicted profit when the treatment indicator is set to 0 according to formula (2). In our applications, we use the zero-inflated negative binomial regression model due to overdispersion in the train dataset. Hurdle models were also tested but showed no improvement compared to the zero-inflated models. The R package pscl is applied. Before estimating the models based on the train data and comparing the results on the test data—as usual in machine learning—reflections on performance evaluation and parameter tuning are necessary. As already discussed in Sect.3, the incremental profit uplift curve and the derived profit Qini coefficient are suitable measures for this purpose. Since only one observation per customer is available in the data (profit if treated or profit if not treated due to the belonging to the treatment or the control group), for calculating uplifts a grouping of customers and comparing average profits of treated and not treated customers in each group is needed. This grouping is based on sorting the customers according to the developed scoring system (starting with the customers where we assume the highest profit uplift) and forming quantiles (usually deciles) of the sorted customers. Basing on these groupings, now, the incremental profit uplift across the quantiles can be plotted (the incremental profit
661 1 3 Profit uplift modeling fordirect marketing campaigns:… uplift curve), and the area between this curve and a curve derived by random sorting (the profit Qini coefficient Q) can be calculated and used for selecting best scoring systems. Figure2 shows the profit Qini coefficients for the four discussed models (OLS, Heckman, and RF as direct models, Zeroinfl as interaction model) when estimated with varying numbers of predictors on the basis of the calibration subsample of the train data and used for predictions on the basis of the validation sample. Note that the profit Qini coefficients reflect the area between the profit Qini curve and the curve for the random model as in Fig.1 and that larger values indicate a better sorting of the customers according to their “observed” profit uplift (calculated via groups of customers with similar uplift predictions). It can be easily seen that the profit Qini coefficients are low with small numbers of predictors as well as with high numbers of predictors. These findings are consistent with the findings of Devriendt etal. 2018), who found in their comparison of binary uplift models that 5–15 predictors typically provide the best results. Against this background, we decided to use 20 predictors in the following for training and testing our models and to compare them with revenue response and uplift as well as profit response models. It can also be seen that overall the two direct regression models regression (OLS and Heckman) performed quite similar in this tuning analysis. In the following, when we concentrate on the 20 principal components as a result of hyperparameter tuning but now analyze the stability of these results in more detail, consequently, we additionally applied the Heckman two-model approach to elaborate further differences. Table5 and Fig.3 already show the results of this extended evaluation of profit uplift modeling approaches: Four profit uplift modeling approaches were applied Fig. 2 Profit Qini coefficients for the validation set (3/7 of the BAUR train set) based on training the profit uplift modeling approaches on the calibration set (4/7 of the BAUR train set)
662 D.Baier, B.Stöcker 1 3 Table 5 Results of the application of profit uplift modeling approaches to the BAUR dataset (20 principal components): 50 random subsamples (6/7) of the train set (50%) were used to calibrate the models and predict the profit uplift in the test set (50%) The profit Qini coefficients were averaged (standard deviations in brackets) Modeling approach Train set Test set Estimation Profit Qini coefficient Profit Qini coefficient Profit uplift Direct OLS 0.495 (0.022) 0.421 (0.013) Two model Heckman 0.228 (0.084) 0.148 (0.097) Direct RF 0.568 (0.026) 0.344 (0.009) Interaction Zeroinfl 0.471 (0.024) 0.402 (0.012) Fig. 3 Application of profit uplift modeling approaches to the BAUR dataset (20 principal components): 50 random subsamples (6/7) of the train set (50%) were used to calibrate the models and predict the profit uplift in the test set (50%). The resulting profit Qini curves (left) and mean profit uplifts (right) were averaged
663 1 3 Profit uplift modeling fordirect marketing campaigns:… to 50 randomly drawn subsamples (6/7) of the train set (50% of the BAUR dataset). The profit Qini coefficients were calculated and averaged (see mean values and standard deviations in Table5 and the mean Qini curve in Fig.3). Then, the estimated profit uplift models were applied to the test data, and—again—the profit Qini coefficients were calculated and averaged (mean values and standard deviations in Table5, mean Qini curve in Fig.3). The results reflect the results of parameter tuning (Please note that since the OLS and Heckman direct model performed similar, therefore in Table5 and Fig.3 the results of the Heckman two model approach are given instead): the direct model with OLS parameter estimation performs best with respect to the holdout test set, followed by the Zeroinfl interaction model, and the RF direct model. The Heckman two-model approach is inferior to these three approaches. However, as Fig.3 demonstrates, OLS, RF, and Zeroinfl provide quite similar results, which is—to some extent – surprising since the modeling assumptions (“normal shape” vs. count data, direct model vs. difference of two predictions based on the interaction model) and the estimation algorithms (one-step vs. two-step estimation) are very different. It should be mentioned that the profit Qini coefficients in Table5 and the Qini curves in Fig.3 are used to select a “best” predictive model and not for deciding on “best” customers for the direct marketing campaign. All customers of the train set and the test set were allocated to the treatment or the control group, and the Qini coefficients and the Qini curves just reflect whether a derived model from the train set is able to correctly predict the uplifts in the test set. When the decision with respect to a best model is made, this best model then is applied to all customers and customers with a predicited positive profit uplift should be included into the direct marketing campaign. However, for these final step, no Qini coefficients or Qini curves can be derived since the necessary balanced distribution of respondents in the tran and test group is not given. The application at least shows that it seems to be possible that—besides already existing binary uplift and revenue uplift models—it is possible to estimate profit uplift models which show clear practical advantages. In the following, we analyze this theoretical superiority using the BAUR dataset by comparing the new approaches with traditional ones. 4.3 Comparison ofrevenue andprofit response anduplift modeling approaches A detailed comparison of the proposed profit uplift modeling approaches to already known revenue response and uplift modeling approaches, but also profit response modeling approaches is used to clarify differences between these approaches. Again, the train set (50%) and the test set (50%) of the BAUR dataset is used for comparisons, 50 subsamples (6/7) of the train set were drawn and used to calibrate the various models under study. For each model and for each subsample of the datasets, scores are predicted for the customers in the data used for training and for the customers in the test set. Please note that in the case of response models, these scores reflect a revenue or profit response (depending on the model), and in the case of uplift models, these scores reflect a revenue or profit uplift. Based on the sorting of the customers according to these predicted (revenue or profit, response or uplift)
664 D.Baier, B.Stöcker 1 3 scores, then, revenue and profit Qini curves, as well as revenue and profit Qini coeffients, can be calculated. These coefficients were averaged across the 50 random subsamples similar as in the previous subsection (Indeed, the profit Qini values for the profit uplift modeling approaches are the same in Tables5 and 6). Table6 as well as Figs.4, 5, and 6 reflect the results of these modeling and prediction endeavors (Please note that Figs.4 and 5 show revenue uplift curves whereas Figs.3 and 6 show profit uplift curves): first, one can see, that altogether 16 modeling approaches were used in this comparison, each applied to 50 subsamples of the train set. The response models were calibrated based on the treated customers in the train set. The aim was to predict the individual revenue or profit of treated customers without taking into account whether the customer would have also bought without being treated. As Fig.4 (for revenue response models) and Fig.6 (for profit response models) demonstrate, these response modeling approaches also convince when the customers of the train and test set should be sorted according to their estimated revenue uplift, but—according to Table6—not when a profit uplift sorting is needed. However, as Table6 clearly demonstrates: the response models are inferior to the uplift models in all cases, i.e. when predicting uplifts is needed and measured via the revenue and the profit Qini coefficients. The same holds when revenue response and uplift are used to predict profit uplifts: Figs.4 and 5, as well as Table6, demonstrate that these models are quite good in predicting revenue responses and uplifts. However, the profit Qini curves (not shown in the Figures) evaluated via the profit Qini coefficients in Table6 show negative values, which means that these models are inferior even to a random model. Finally, Fig.6 also shows that for profit uplift prediction, the application of a profit response model is not enough. To summarize: the extensive comparison of various response and revenue uplift models applied to the BAUR dataset reflects promising results for the usefulness of the new profit uplift modeling approaches. Even an application with a simple (onestage) regression or a random forest estimation algorithm applied to transformed profit data outperforms these traditional models. In the following, we investigate whether this superiority can also be found when a well-known and publicly available dataset is used, which has often been the basis for introducing new uplift modeling approaches. 5 Application totheHillstrom dataset In order to demonstrate that the new profit uplift modeling approaches are applicable and superior to traditional response and uplift modeling approaches, also a standard dataset from the uplift modeling literature is analyzed, the Hillstrom dataset (Radcliffe 2008). This dataset was made available by Kevin Hillstrom through his MineThatData blog and contains a sample of 64,000 customers which had been divided up into three nearly equally sized subsamples, two of them contacted via two direct marketing campaigns and one not contacted, serving as a control group (see the similar usage of this dataset, e.g., in Rudaś and Jaroszewicz 2018). Table7 summarizes the descriptive uplift statistics of this dataset, where the two treated
665 1 3 Profit uplift modeling fordirect marketing campaigns:… Table 6 Results of the application of revenue and profit response and uplift modeling approaches to the BAUR dataset (20 principal components): 50 random subsamples (6/7) of the train set (50%) were used to calibrate the models and predict the revenue and profit uplift in the test set (50%) The revenue and profit Qini coefficients were averaged (standard deviations in brackets) Modeling approach Train set Test set Estimation Revenue Qini coefficient Profit Qini coefficient Revenue Qini coefficient Profit Qini coefficient Revenue response Direct OLS 2.612 (0.099) −0.329 (0.019) 2.370 (0.026) −0.318 (0.006) Heckman 2.424 (0.096) −0.302 (0.018) 2.133 (0.000) −0.295 (0.000) RF 2.920 (0.102) −0.296 (0.018) 2.355 (0.018) −0.317 (0.004) Zeroinfl 2.575 (0.103) −0.300 (0.020) 2.400 (0.017) −0.278 (0.004) Revenue uplift Direct OLS 2.744 (0.088) −0.178 (0.022) 2.531 (0.045) −0.162 (0.020) Two model Heckman 2.648 (0.135) −0.106 (0.058) 2.227 (0.102) −0.139 (0.055) Direct RF 3.785 (0.126) 0.026 (0.032) 2.438 (0.036) −0.234 (0.012) Interaction Zeroinfl 2.526 (0.109) −0.127 (0.028) 2.370 (0.079) −0.118 (0.019) Profit response Direct OLS 2.612 (0.099) −0.329 (0.019) 2.370 (0.026) −0.318 (0.006) Heckman 2.504 (0.119) −0.296 (0.020) 2.210 (0.090) −0.286 (0.006) RF 2.920 (0.102) −0.296 (0.018) 2.355 (0.019) −0.317 (0.004) Zeroinfl 2.575 (0.103) −0.300 (0.020) 2.400 (0.017) −0.278 (0.004) Profit uplift Direct OLS −0.118 (0.150) 0.495 (0.022) −0.427 (0.101) 0.421 (0.013) Two model Heckman 1.118 (0.391) 0.228 (0.084) 0.655 (0.281) 0.148 (0.097) Direct RF −0.778 (0.202) 0.568 (0.026) −1.990 (0.069) 0.344 (0.009) Interaction Zeroinfl 0.306 (0.164) 0.471 (0.024) −0.025 (0.109) 0.402 (0.012)
666 D.Baier, B.Stöcker 1 3 subsamples are merged. As can easily be seen, the conversion rate is much lower as in the BAUR dataset (on average 1.07% in the treatment group) but nevertheless shows a conversion rate uplift compared to the control group (on average, an uplift of 0.50%). The revenue uplift per customer is 0.60$, but this uplift seems to be arising solely from the conversion rate uplift since the average revenue spend by a purchaser in the treatment group (117.00$) is only slightly higher than in the control group (114.00$). Again, as in the BAUR dataset, we assume that the campaign offers a 20% discount and that the margin for the retailer is 30%. With these assumptions (not part of the original communication of the dataset, just an assumption to be able to analyze the dataset with our profit uplift modeling approaches), the overall profit uplift per customer is negative (−0.07$). So, again we have to develop a scoring system that helps to restrict the direct marketing campaign to customers with a positive profit uplift prediction. Fig. 4 Application of revenue response modeling approaches to the BAUR dataset (20 principal components): 50 random subsamples (6/7) of the train set (50%) were used to calibrate the models and predict the revenue uplift in the test set (50%). The resulting revenue Qini curves (left) and mean revenue uplifts (right) were averaged
667 1 3 Profit uplift modeling fordirect marketing campaigns:… The original dataset also contains potential predictors for this scoring system, as given in Table 8. The original eight potential predictors (in Table 8 described as variable categories) were scaled nominally (e.g., history_segment with 7 values or channel with three values) or metrically (e.g., recency or history). For our further analysis with the three models, we dummy-coded the nominally scaled potential predictors and so received in total 25 metrically scaled variables (see Table8). As in Sect.4, the customers were randomly partitioned into a train set (~ 70% or 44,800 customers) and a holdout test set (~ 30% or 19,200 customers), and the train set was preprocessed by setting means to zero, setting standard deviations to 1, and applying a Box–Cox-transformation to transform skew distributed variables into “normal shape”. The same preprocessing was applied to the test set, using the transformation parameters derived from the train set. Then, five models, similar as in Sect.4, were applied: 50 subsamples (6/7) of the train set were Fig. 5 Application of revenue uplift modeling approaches to the BAUR dataset (20 principal components): 50 random subsamples (6/7) of the train set (50%) were used to calibrate the models and predict the revenue uplift in the test set (50%). The resulting revenue Qini curves (left) and mean revenue uplifts (right) were averaged
668 D.Baier, B.Stöcker 1 3 drawn randomly, the models were calibrated, and the Qini curves and Qini coefficients were averaged. Figure7 and Table9 reflect the results of this modeling and prediction task. It should be mentioned that we only use a subsample of the models applied in Sect.4 but the selection contains the “best” models from this comparison (e.g. especially the three “winners” OLS and RF direct models and Zeroinfl interaction model). One can easily see that the three “best” profit uplift modeling approaches (onestage OLS and RF as well as two-stage Zeroinfl), again, show similar results with random forest providing the best performance. But it should be mentioned that—maybe due to the few purchasers in the dataset with a high concentration of revenues and profits from few purchasers—the modeling leads to a worse performance compared to the application of the BAUR dataset. This problem of the Hillstrom dataset when it comes to modeling continuous outcomes has also been Fig. 6 Application of profit response modeling approaches to the BAUR dataset (20 principal components): 50 random subsamples (6/7) of the train set (50%) were used to calibrate the models and predict the profit uplift in the test set (50%). The resulting profit Qini curves (left) and mean profit uplifts (right) were averaged