Full text
ORIGINAL ARTICLE Application of choice models in tourism recommender systems Ameed Almomani 1 | Paula Saavedra 1 | Pablo Barreiro 1 | Roi Durán 1 | Rosa Crujeiras 2 | María Loureiro 3 | Eduardo Sánchez 1 1 CITIUS, University of Santiago de Compostela, Santiago, Spain 2 School of Mathematics, University of Santiago de Compostela, Santiago, Spain 3 School of Business, University of Santiago de Compostela, Santiago, Spain Correspondence Eduardo Sánchez, CITIUS, University of Santiago de Compostela, Santiago, Spain. Email: eduardo.sanchez.vil[email protected] Funding information Consellería de Cultura, Educaci on e Ordenaci on Universitaria, Xunta de Galicia, Grant/Award Number: ED431G/08; EMALCSA, Grant/Award Number: CSC-14-13; Ministerio de Ciencia e Innovaci on, Grant/Award Number: TIN2014-56633-C3-1-R; Ministerio de Economía, Industria y Competitividad, Gobierno de España, Grant/Award Number: MTM2013-41383P; European Regional Development Fund Abstract Choice models (CM) are proposed in the field of tourism recommender systems (TRS) with the aim of providing algorithms with both a theoretical understanding of tourist's motivations and a certain degree of transparency. The goal of this work is to overcome some of the limitations of current state-of-art algorithms used in TRSs by providing: (1) accurate preferences, which are learnt from user choices rather than from ratings, and (2) interpretable coefficients, which are achieved by means of the set of estimated parameters of CM. The study was carried out with a gastronomic data set generated in an ecological experiment in the tourism domain. The performance of CM has been compared with a set of baseline algorithms (rating-based and ensembles) by using two evaluation metrics: precision and DCG. The CM outperformed the baseline algorithms when the size of the choice set was limited. The findings suggest that CM may provide an optimal trade-off between theoretical soundness, interpretability and performance in the field of TRS. KEYWORDS artificial intelligence, choice models, ensembles, knowledge engineering, recommender systems, tourism 1|INTRODUCTION Tourism is a strategic business domain which manages the movement of people to destinations outside their usual environment and plays a key role in the economic and social development of many countries. In the last decade, the digitalization of the sector has been crucial to reach new markets and offer new experiences to tourists. In this field, intelligent or smart tourist recommender systems (TRS) have become mainstream techniques for suggesting personalized plans, products and information to end users. The field of Machine Learning has played an important role in providing a number of recommendation techniques that work at the back-end of TRSs (Borràs et al., 2014; Hamid et al., 2021; Kzaz et al., 2018). The most relevant approaches are: content, collaborative, context and ensemble-based models. Content-based algorithms work both with item attributes and past user experiences to learn a profile of preferences for each decision-maker (Burke et al., 2011; Harman, 1995). Tourist profiles are built based on demographic or location-based information which is used to estimate the relevancy of point of interests (Santos et al., 2019). Collaborative-based algorithms, on the other hand, require users to rate items, which are later used to build memory and model-based approaches to predict the rating of any new user-item interaction (Burke et al., 2011; Resnick et al., 1994). Since its arrival, the collaborative approach has been widely adopted as a way of removing the need to manage specific domain information, and ratings have become the key data required to fuel rating-based algorithms. While content and collaborative approaches are based on user-item interaction, context-based techniques include contextual information that informs about the specific Received: 5 April 2022 Revised: 7 October 2022 Accepted: 15 October 2022 DOI: 10.1111/exsy.13177 This is an open access article under the terms of the Creative Commons Attribution-NonCommercial-NoDerivs License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non-commercial and no modifications or adaptations are made. © 2022 The Authors. Expert Systems published by John Wiley & Sons Ltd. Expert Systems. 2023;40:e13177. wileyonlinelibrary.com/journal/exsy 1of18 https://doi.org/10.1111/exsy.13177
circumstances of such interaction (Adomavicius & Tuzhilin, 2011). Objects and points of interest can be used as contextual information to suggest optimal plans for tourist (Le & Pishva, 2016). Nowadays, after the impact of the Netflix prize, ensemble models have become very popular to develop recommender systems. Ensemble learning is a paradigm that aggregates instances of weak algorithms (aka learners) to produce more accurate predictions than those provided by single learners (Friedman et al., 2009; Polikar, 2006). The most popular methods proposed for efficient aggregation of learners (Friedman et al., 2009; Polikar, 2006) are bagging, boosting and random forest. A hybrid ensemble, which combines learners of different nature, has been successfully applied to recommend tourist routes based on location-tagged data (Wan et al., 2018). The main drawbacks of the techniques applied so far in the development of TRSs are: the lack of a theoretical background to understand the underlying motivational factors conditioning the tourist decision-making, and their interpretability to explain the recommendations. The study and identification of tourist's motivation is a central element of the so called push-pull models in tourism (Crompton, 1979; Dann, 1976; Pestana et al., 2020). They are based on the notion that tourists make choices according to a set of needs and motivations that push them to travel, while tourist items have a set of desirable characteristics that attract them. So, the first pillar to develop a sound TRSs should be to understand and learn those push motivations that may explain both tourist's preferences and behaviours. The machine learning approaches apply different shortcuts in order to solve this problem. All rating-based models learn preferences from ratings considering a strong relationship between them: in memorybased approaches, it is assumed that decision-makers with similar ratings will have similar tastes; and in model-based techniques, ratings are assumed to be the result of a matching between latent factors in an item and the decision-maker's preferences about those factors. The problem is the absence of experimental evidence supporting these assumptions, which makes the accuracy of the learnt tastes/preferences unclear. Furthermore, ensemble-based solutions, while successful in terms of accurate predictions, come at a cost of complexity and an opaque nature that makes the recommendations difficult to explain. Our work focuses on providing a theoretical background to the algorithms behind tourism recommender systems to alleviate these problems. The contribution to the field of TRSs can be summarized as follows: •The application of choice models (CM) as a tool to learn tourist's preferences with a sound methodology. •The exploration of the potential of CM by comparing them with algorithms used in TRSs: (1) advanced rating-based algorithms and (2) ensemble strategies. The paper is organized in the following way. In the Related Work section, an overview of CM and other choice-based strategies are reviewed. In the Background section, the recommendation problem is described as a choice problem and the CM are presented. In the Methods section, the experiments, datasets and algorithms are described. In the Results section, the fitting as well as the performance evaluation of the algorithms are presented. Finally, in the Discussion section we comment on the results and highlight the major contributions of the paper. 2|RELATED WORK 2.1 |Choice models Chaptini proposed the first application of discrete choice models (CM) in the field of recommender systems (Chaptini, 2005). The goal was to provide personalized course recommendations for MIT students by means of a generalized mixed logit model fed with survey data. Some years later, Polydoropoulou and Lambrou continued the utilization of CM to recommend courses for seafarers and employees of the shipping industry (Polydoropoulou & Lambrou, 2012). The models were estimated with data gathered from questionnaires. The novelty was centred on the application of a Bayesian approach to update the estimated coefficients, which allows for different prior distributions to characterize individual preferences. Amore sophisticated model, a multi-level nested multinomial logit one, was proposed by Jiang et al. in the quest of achieving both relevancy and diversity in the recommendation process (Jiang et al., 2014). Recently, the group of Ben-Akiva at MIT have explored the potential of CM in app-based recommender systems (Danaf et al., 2019), a setting in which the attributes of the alternatives may vary over time and therefore user's preferences need to be continuously updated. The updating method has been tested in the field of transportation with real choices of users collected in Switzerland. In the field of tourism, our preliminary work has revealed the potential value of CM for gastronomic recommendations by comparing their performance with that of basic rating-based algorithms (Saavedra et al., 2016). Following a similar line of thought, Mottini and Leheritier analysed how CM can be used in the air travel industry to have a better understanding of flight choices (Mottini et al., 2018). 2.2 |CM and the quality of data Two types of information can be gathered from users: explicit and implicit. Explicit ratings obtained directly from users are the most common information used in recommender systems. However, it has been pointed out that ratings have some important drawbacks (Claypool et al., 2001): 2of18 ALMOMANI ET AL. 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
(1) the input of these opinions can alter normal pattern of browsing and reading, and (2) users may stop entering this data if they do not see the benefit. As a result, the ratings may not be a reliable source of data for the recommending process. The solution is the use of implicit information, that is, the type of data obtained using an indirect method without direct interrogation of the user. Actions such as mouse clicks, mouse movement, browsing times and user's choices can be recorded and used to derive user's interest and preferences (Peska & Vojtas, 2017). Choices are the key data proposed in this paper to feed both the CM and the TRSs. 2.3 |CM and patterns of human decision-making Understanding how tourists make choices is crucial to develop efficient TRSs. On this regard, the ASPECT model comes handy as it describes six different decision-making strategies or patterns that humans may follow when facing a choice problem (Jameson et al., 2015): attribute-based, consequence-based, experience-based, socially-based, policy-based and trial-and-error-based choice. A relationship between these patterns and state-of-art recommendation algorithms can be found: content-based algorithms follow the ideas behind the attribute-based pattern while collaborative-based algorithms could be related with the socially-based one. The CM utilized in this work may be considered as a formal implementation of the attribute-based choice pattern. 3|BACKGROUND ON CHOICE MODELS 3.1 |Recommendation as a choice problem The recommendation problem can be approached in different ways by viewing it as the problem of predicting user's choices in any particular context. Under this perspective, the Rational Choice Theory can be considered the classic paradigm used to explain the choices made by rational agents (Sen, 1990). This theory assumes that any decision-maker will solve the decision-making problem by applying the following rule: CR A,ðÞ¼a0A a0a,8aA no ,ð1Þ where, CR represents ‘choice rule’,Ais the choice set, the set of alternatives considered for the decision maker at the time of choice, and the operator represents the relationship ‘preferred to’, or at least ‘preferred’. The chosen alternative will therefore be that for which the decisionmaker shows the greatest preference. In order to build a predictive model on the basis of this rule, the researcher must replace the qualitative preference operator with a quantitative one that will enable numerical comparison between the benefit of each alternative. Utility theory comes to the rescue to solve this issue. One of its axioms states that it is possible to define a utility function such that, ab,Ua ðÞ ≥Ub ðÞ :ð2Þ Therefore the choice rule in Equation (1) becomes: CR A,≥ðÞ¼a0A Ua 0 ðÞ≥UaðÞ,8aA no :ð3Þ This rule is mathematically equivalent to the formulation of the general recommendation problem (Adomavicius & Tuzhilin, 2005), which is described in terms of a maximization problem: a0¼argmax aA UaðÞ, CR A,≥ ðÞ ¼a0A Ua 0 ðÞ ≥Ua ðÞ ,8aA no : ð4Þ As the recommendation problem can be understood as a choice prediction problem, the powerful models and techniques developed in the latter field can be applied to generate recommendations. ALMOMANI ET AL.3of18 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
3.2 |Choice models with random utility The choice rule represents how decision-makers reach their decisions. However, in the real world, researchers do not have access to all of the information that decision-makers may handle to estimate the utilities. For a specific user c n , the researcher only knows some attributes of the alternatives, labelled x j , for all a j alternatives with j{1, …,J}. Therefore, the predicted utility can be decomposed as follows: Uc n,aj ¼Vnj þϵnj,ð5Þ where, V nj =V(x j ) is the representative utility, which can be estimated on the basis of the observed factors, and ϵ nj captures the unknown factors that cannot be observed by the researcher. This decomposition is fully general, as ϵ nj is defined simply as the difference between the true utility U nj and the representative utility V nj . The uncertainty about ϵ nj is handled as a random variable, and the researcher must make further assumptions about its probability distribution. The models derived under these assumptions are called random utility models (RUM) (McF, 1973). From the researcher's perspective, the choice rule of Equation (3) for a decision-maker c n , which is deterministic from the decision-maker's perspective, becomes probabilistic in the following way: CR A,≥ðÞ¼aiA Pni ≥Pnj,8ajA no ,ð6Þ and the probability Pni is estimated by considering the decomposition formulated in Equation (4): Pni Uc n,ai ðÞ >Uc n,aj for all j≠i ¼Pni ϵnj ϵni <Vni Vnjfor all j≠i :ð7Þ If the joint density of ϵ n =(ϵ n1 ,…,ϵ nJ ) is denoted by f, the cumulative probability can be rewritten as follows: Pni ¼ðϵ ϵnj ϵni <Vni Vnj for all j≠i fϵn ðÞdϵn,ð8Þ where, is the indicator function, equalling 1 when the term in parentheses is true and 0 otherwise. 3.3 |Standard and mixed logit models Different models are derived depending on the density chosen, that is, depending on the evidence or assumptions about the distribution of the unobserved portion of utility. The simplest and most widely adopted choice model is the standard logit model (McF, 1973), which is obtained under the assumption that each unobserved portion of utility ϵ nj is distributed independently and identically. In this case, fdenotes the density for Gumbel distribution and the integral 8 takes a closed form with the following solution: Pni ¼eVni PjeVnj :ð9Þ This model estimates the probability Pni as the ratio between the relevancy of the item a i for user c n , estimated by the eVni term, and the aggregated relevancy of all items a j in the choice set. This set is the collection of items that the user considers/analyzes at the time of choice. Typically it is a reduced number of items that were filtered by the user by considering different constraints (price, distance, knowledge of the user, etc.). The values of the probability Pni depend on the representative utilities. As V ni increases, reflecting a higher match between the observed attributes of the alternative and the preferences of the decision-maker, with V nj for all j≠iheld constant, Pni approaches the value one. Pni approaches zero when V ni decreases, as the exponential in the numerator approaches zero as V ni approaches ∞. The representative utility is usually specified as linear in the set of alternative attributes: V nj =β nj x j , where x j is a vector including, as before, the observed attribute's values of the alternative a j , and β nj denotes the model coefficients vector describing the preferences of decision-maker c n for the attributes of the alternatives a j . The preferences β nj (model coefficients) are estimated by fitting Equation (9) to a data set of choices. The choice set must verify three properties. It must be finite, exhaustive (the decision-maker always chooses one of the alternatives) and mutually exclusive (the choice of one alternative necessarily implies not choosing any of the other ones). 4of18 ALMOMANI ET AL. 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
The standard logit model cannot represent differences in tastes that are not related to observed characteristics (Train, 2009). Therefore, if taste variation is modelled as partly random, a logit model with random parameters should be considered instead. Thus, βis now a vector of random coefficients that vary across decision-makers in the population with density g. This density is a function of parameters θthat represent, in the Gaussian case, the mean and covariance of the random coefficient in the population. The choice probabilities can be written as follows: Pni ¼ðeVni βðÞ P j eVnj βðÞ 0 B @1 C AgβjθðÞdβ:ð10Þ As the previous integral does not adopt a closed form, it must be evaluated numerically. Once the researcher specifies a distribution gfor the coefficients, the parameters θmaximizing the simulated log-likelihood must be estimated through simulation. The Rdraws of the coefficients are then taken from gand the logit probabilities are computed for each draw. The unconditional probability in Equation (10), which is the expected value of the conditional probabilities, is estimated as the average of the Rprobabilities determined previously. 3.4 |Required data In order to fit a choice-based model, we need a sufficient number of choices taken by the decision-maker. For each choice, the following data is required: •The vector x i for the chosen alternative. •The vectors x j for all alternatives a j in the choice set. 4|METHODS The methods were chosen to compare the performance of CM against rating-based models and popular ensemble strategies. The analysis was carried out with a gastronomic data set generated in an ecological experiment in the tourism domain. The design is described in Section 4.1 and the data set in Section 4.2. The details of the CM considered in this study are presented in Section 4.3. The baseline algorithms (rating-based and ensembles) chosen to compare our models are introduced in Section 4.4. The evaluation criteria used to estimate the performance of each algorithm are included in Section 4.5. Finally, software and implementation details are provided in Section 4.6. 4.1 |Experiment We designed an ecological experiment under the scope of the RECTUR project. The chosen setting was the fourth edition (in 2011) of the Santiago(é)Tapas contest, a gastronomic event that takes place every year in the city of Santiago de Compostela. For the event, 56 local restaurants proposed and elaborated up to three tapas that were sold at a fixed price. A total of 5517 participants, including local, Spanish and international users, tasted the available tapas over a period of 2 weeks. A TapasPassport was made available to all participants and included the following official information: (i) the contest guidelines, (ii) restaurant location, and (iii) the tapas offered at each restaurant. After consuming the tapas, participants evaluated their experience by providing a vote with two ratings (Figure 1): (i) a rating for the tapas, and (ii) a rating of the overall experience (service, place atmosphere, etc.). It is important to point out that the experiment was carried out in a real setting rather than a laboratory setting. Thus, the restaurants were free to offer whatever type of tapas they wished, and the participants made their own decisions about which tapas to try. It can therefore be assumed that the data set will include some sampling bias that may have some impact on the model predictions. 4.2 |RECTUR datasets The data collected in the experiment were used to build two datasets: the choice and the rating dataset. The choice dataset consisted on choice observations, where each observation included the vector x i and the vectors x j containing the attribute values of the chosen tapa a i as well as the tapas a j of the choice set. To describe a tapa, the following attributes were considered (Table 1): type and character. Traditional tapas were ALMOMANI ET AL.5of18 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
created following well-known, popular recipes, while daring tapas were new and creative. In terms of data preparation, type and character attributes were transformed into eight dichotomous or binary variables associated with each value. Suppose the observation of a decision-maker located in the old area of the city choosing tapa t100, which is of a meat type and has a daring character. In this case, the chosen tapa was codified as follows: (1) ‘meat’variable set to 1, (2) all other type variables set to 0, (3) ‘daring’variable set to 1, and (4) ‘traditional’variable set to 0. The choice dataset codified this way was used to fit the choice-based models. On the other hand, the rating dataset stored a collection of ratings, an attribute of the tapa-user interaction, to gather the user's satisfaction with the tapa. The rating dataset was applied to train the baseline models. 4.3 |Choice models: standard and mixed logit models The standard logit model as well as the mixed logit model, assuming Gaussian distribution on the coefficients, were chosen as basic representatives of the family of random utility choice-based models. Application of the mixed logit model was justified as we found evidence of taste variations among decision-makers on the basis of both personal and contextual factors (Ismoilov, 2017). Although a large number of users tasted more than one tapa, the number of choices per user were not enough to fit a choice model per user. Constrained by this limitation, we decided to define three choice problems, each one corresponding to each area of the city. Each choice problem therefore aggregated the observations of all choices on each area of the city and assumed an unique choice set for all users in that area. Both the standard and mixed logit models were estimated for the three problems. FIGURE 1 RECTUR experiment. Images of votes, participating locals, and TapasPassport. The experiment was carried out in Santiago de Compostela during the celebration of a real contest of tapas TABLE 1 Tapa and tapa-user attributes Tapa attribute Values ID t1totn Type Cheese, Egg, Fish, Meat, Vegetable, Shellfish, Sweet, and Other Character Traditional or daring Tapa-user attribute Values Rating 0–5 6of18 ALMOMANI ET AL. 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
4.4 |Baselines The following types of baseline models were chosen: (1) basic rating-based collaborative filtering algorithms, (2) advanced rating-based collaborative algorithm, (3) single decision-trees, and (4) tree-based ensemble strategies. The choice of tree-based from among other types of ensembles is explained by the fact that trees produce meaningful predictions (Ali et al., 2015; Quinlan, 1986), and they thus become a natural alternative to choice-based models to overcome the interpretability problem. Moreover, tree-based methods have proved useful for building recommender systems in different areas outperforming other approaches (Utku et al., 2015). However, the accuracy of prediction may suffer from both the size of the available data set (Bar et al., 2013; Ghimire et al., 2012) and the number of features (Lavanya & Rani, 2012). Our previous studies have demonstrated the superior performance of tree-based ensembles relative to single decision-trees, as well as their dependency on the number of available features (Almomani et al., 2017). 4.4.1 | Basic rating-based collaborative filtering (CF) Two basic rating-based CF models were used: user-based collaborative filtering (CF-UB) and matrix factorization (CF-MF). CF-UB assumes that individuals with similar preferences will rate items in a similar way. Thus, missing ratings for a specific user c n can be predicted by finding a neighbourhood N(n) of similar users and aggregating their ratings to calculate the corresponding prediction. The concept of similarity between users is used to define the neighbourhood given all users within a similarity threshold. In this study, the cosine similarity measure was considered, and jN(n)jwas set at 25. For an item iand an individual c n , the ratings predicted, b rni, can be expressed as follows: b rni ¼1 jNnðÞjX jNnðÞ rji,ð11Þ where, jj denotes the cardinality of N(n). CF-MF, on the other hand, characterizes both items and users by vectors of factors inferred from item rating patterns. For a given item iand a user c n , the vector q i measures the extent to which the item possesses those factors and the vector p n , the extent of interest the user has in items that score highly on the corresponding factors. The dot product qT ipncaptures the user's interest in the item's characteristics. This approximates user c n 's rating of item i,r ni , leading to the following estimate: b rni ¼qT ipn:ð12Þ Therefore, the challenge is to compute the mapping of each item and user to vectors q i and p n . Here, singular value decomposition will be applied to factoring the user-item rating matrix, which may be sparse. In order to learn the factor vectors (p n and q i ), the regularized squared error on the set of known ratings is minimized: minq,pX u,iðÞK rni qT ipn 2þλqi kk 2þpn kk 2 ,ð13Þ where, Kis the set of the (c n ,i) pairs for which r ni is known, kk is the Euclidean norm and λdenotes a constant controlling the extent of regularization. In this work, λ=1.5. 4.4.2 | Advanced rating-based collaborative filtering (CF) As a more complex model of this family, we resorted to CF-SVD++, an extension to CF-MF in which the effect of implicit information is included in the minimization rule. The difference here is that the prediction rule considers the fact of a user rating of an item as an additional indication of preference. Therefore, the vector representing the user's interest becomes (Koren, 2008): pnþjNnðÞj 1 2X jNn ðÞ yj:ð14Þ ALMOMANI ET AL.7of18 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
4.4.3 | Single decision trees The idea of decision trees is to build a tree structure with nodes representing the features or attributes and leaves indicating the corresponding values of the attributes. The trees can be used for classification or for regression depending on the nature of the predicted outcome (Breiman et al., 1984). In this study, we used regression trees as the predicted variables (i.e., ratings) are numerical. The tree is constructed through binary recursive partitioning, an iteration process that splits the features into branches. The process continues by splitting each partition into a minimum number of nodes. For the recursive binary splitting, both the splitting variable X i and a split point z are considered. The splitting at the split point is therefore described as follows: R1i,zðÞ¼XjXi≥z fg and R2i,zðÞ¼XjXi<z fg :ð15Þ A tree is formally described as follows: TX,ΘðÞ¼ X J j¼1 γjIXRj ,ð16Þ where a γ j parameter is assigned to each terminal node, and Θ={R j ,γ j }. The prediction will be the mean of the outcome predictions (i.e., rating predictions in this study) in the region or terminal node. Figure 2shows an example of the regression tree learnt for user number 1377 (u1377) and 44 tapas. FIGURE 2 Regression tree for user number 1377 learnt from 44 consumed tapas 8of18 ALMOMANI ET AL. 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
4.4.4 | Tree-based ensemble strategies Ensembles are aggregations of simple learners, such as trees, and the final prediction is estimated by combining the outcomes. The three ensemble methods used in this study are described below: •Boosting builds a tree by using an iterative procedure. This means each tree depends and improves its performance on the basis of the prior trees. The prediction is estimated as follows: fmxðÞ¼ X M m¼1 TX,θm ðÞ,ð17Þ where, Mis the number of trees. •Bagging builds different trees on Mdifferent bootstrapped training data set. All trees are fully grown, indicating that a search over all features is carried out at each node in order to find the feature that best splits the data at that node. The final prediction is the average of each single tree estimation b fbxðÞ: b fbag xðÞ¼ 1 MX M m¼1b fbxðÞ:ð18Þ •Random Forests (RF) is a particular case of Bagging. The main difference is that at each candidate split in the learning process, a random sample of the predictors or features is chosen among all the predictors or features. The goal is to build a large collection of uncorrelated trees. The prediction in a regression problem is estimated as follows: b frf xðÞ¼ 1 MX M m¼1b fbxðÞ:ð19Þ 4.5 |Evaluation The performance of all models in the three areas of the city was analysed by applying random sub-sampling and leave-one-out cross validation to the RECTUR data set. For validation of random sub-sampling, 100 iterations were considered using 25% of randomly selected individuals for testing and the other 75% for training. For each decision-maker in the test data and for each recommendation method, prediction error measures were then estimated. The procedure for leave-one-out cross validation is similar, but the test set includes only one decision-maker per iteration. Two metrics were applied in order to evaluate the performance of choice-based and rating-based algorithms: Precision and Discounted Cumulative Gain (DCG). For each tapas item included in the choice set, either its rating or its choice probability was predicted. Thereafter, the tapas were ranked and only the item with highest value was considered the predicted choice and therefore recommended (Top-1 scheme). The Precision measure was estimated as the fraction of correct recommendations to total recommendations after comparing the predicted choices with the real ones (Salton & McGill, 1986): Precision ¼Correct recommendations Total Recommendations :ð20Þ Discounted Cumulative Gain (DCG) was chosen as a measure of ranking quality to capture the distance between the true choice and the predicted choice (Järvelin & Kekäläinen, 2002). This is defined as follows: DCGp¼X p i¼1 reli log2iþ1 ðÞ ,ð21Þ ALMOMANI ET AL.9of18 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
problem the models deal with quite large choice sets. This is so as we assumed the user may consider to taste any tapa available on the region she would be located. Therefore, all the tapas available on each of the three areas of the city should belong to the choice set of that area. In fact, by making this assumption, the choice set coincides with the recommendation set. In general, the choice set should be much smaller than the ones used in our paper, so we believe that the size of the choice set will not be an issue when applied to other problems. •Prediction/Recommendation task. At this stage, the recommendation set is used to estimate the choice probability per each alternative. In this stage, the number of alternatives in the recommendation set is not a real problem, as it only impacts on the computational cost of running Nprobability estimations. In a real e-commerce scenario where timely recommendations are required and Nmay be extremely large, various solutions could be implemented, such as reducing the recommendation set (viewed alternatives, purchased alternatives by similar users, etc.) and/or periodically updating the probability estimations beforehand. 6.4 |Conclusion CM seem to provide an optimal trade-off between theoretical soundness, interpretability and performance in the field of tourism recommender systems. For future work, we plan to go a step further in terms of uncovering motivational drivers of behaviour. For instance, we wonder about the effect of cultural factors in the mindset of tourists and how they affect their choices. We are also interested on analysing different choice problems in the field of tourism in order to prove the efficiency of CM in a broader scope. In summary, CM pave the way to the application of sound decision-making models in the field of tourism recommender systems, and open the door of incorporating new and powerful motivational factors to develop more efficient algorithms. ACKNOWLEDGEMENTS This research was sponsored by EMALCSA/Coruña Smart City under grant CSC-14-13, the Ministry of Science and Innovation of Spain under grant TIN2014-56633-C3-1-R, the Ministry of Economy and Competitiveness of Spain under grant MTM2013-41383P, the Consellería de Cultura, Educaci on e Ordenaci on Universitaria (accreditation 2016-2019, ED431G/08), and the European Regional Development Fund (ERDF). CONFLICT OF INTEREST The authors declare no potential conflict of interest. DATA AVAILABILITY STATEMENT The data that support the findings of this study are available from the corresponding author upon reasonable request. REFERENCES Adomavicius, G., & Tuzhilin, A. (2005). Toward the next generation of recommender systems: A survey of the state-of-the-art and possible extensions. IEEE Transactions on Knowledge and Data Engineering,17(6), 734–749. Adomavicius, G., & Tuzhilin, A. (2011). Context-aware recommender systems. In Recommender systems handbook (pp. 217–253). Springer. Ali, S., Tirumala, S. S., & Sarrafzadeh, A. (2015). Ensemble learning methods for decision making: Status and future prospects. In International conference on machine learning and cybernetics (ICMLC) (pp. 211–216). IEEE. Almomani, A., Saavedra, P., & Sánchez, E. (2017). Ensembles of decision trees for recommending touristic items. In International work-conference on the interplay between natural and artificial computation (pp. 510–519). Springer. Bar, A., Rokach, L., Shani, G., Shapira, B., & Schclar, A. (2013). Improving simple collaborative filtering models using ensemble methods. In International workshop on multiple classifier systems (pp. 1–12). Springer. Borràs, J., Moreno, A., & Valls, A. (2014). Intelligent tourism recommender systems: A survey. Expert Systems with Applications,41(16), 7370–7389. Breiman, L., Friedman, J., Olshen, R., & Stone, C. (1984). Classification and regression trees. Wadsworth and Brooks. Burke, R., Felfernig, A., & Goker, M. (2011). Recommender systems: An overview. AI Magazine,32(3), 13–18. Chaptini, B. H. (2005). Use of discrete choice models with recommender systems [PhD thesis]. Massachusetts Institute of Technology. Claypool, M., Le, P., Wased, M., & Brown, D. (2001). Implicit interest indicators. In Proceedings of the 6th international conference on intelligent user interfaces (pp. 33–40). Croissant, Y. (2012). Estimation of multinomial logit models in r: The mlogit packages. R package version 02-2. Crompton, J. L. (1979). Motivations for pleasure vacation. Annals of Tourism Research,6(4), 408–424. Danaf, M., Becker, F., Song, X., Atasoy, B., & Ben-Akiva, M. (2019). Online discrete choice models: Applications in personalized recommendations. Decision Support Systems,119,35–45. Dann, G. (1976). The holiday was simply fantastic. The Tourist Review,31(3), 19–23. Friedman, J., Hastie, T., & Tibshirani, R. (2009). The elements of statistical learning: Data mining, inference, and prediction. Springer-Verlag. Ghimire, B., Rogan, J., Galiano, V. R., Panday, P., & Neeti, N. (2012). An evaluation of bagging, boosting, and random forests for land-cover classification in cape cod, Massachusetts, USA. GIScience & Remote Sensing,49(5), 623–643. Hamid, R. A., Albahri, A. S., Alwan, J. K., Al-Qaysi, Z., Albahri, O. S., Zaidan, A., Alnoor, A., Alamoodi, A. H., & Zaidan, B. (2021). How smart is e-tourism? A systematic review of smart tourism recommendation system applying data management. Computer Science Review,39(100), 337. 16 of 18 ALMOMANI ET AL. 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Harman, D. (1995). Overview of the 3rd text retrieval conference (trec3). In The 3rd text REtrieval conference (TREC-3) (pp. 1–20). Department of Commerce, National Institute of Standards and Technology. Ismoilov, J. (2017). Stated and revealed preferences on gastronomic tourism in Santiago de Compostela [PhD thesis]. University of Santiago de Compostela. Jameson, A., Willemsen, M. C., Felfernig, A., de Gemmis, M., Lops, P., Semeraro, G., & Chen, L. (2015). Human decision making and recommender systems. In Recommender systems handbook (pp. 611–648). Springer. Järvelin, K., & Kekäläinen, J. (2002). Cumulated gain-based evaluation of ir techniques. ACM Transactions on Information Systems (TOIS),20(4), 422–446. Jiang, H., Qi, X., & Sun, H. (2014). Choice-based recommender systems: A unified approach to achieving relevancy and diversity. Operations Research,62(5), 973–993. Koren, Y. (2008). Factorization meets the neighborhood: A multifaceted collaborative filtering model. In Proceedings of the 14th ACM SIGKDD international conference on knowledge discovery and data mining (pp. 426–434). ACM. Kzaz, L., Dakhchoune, D., Dahab, D., Park, D., Kim, H., Carrer-Neto, W., Hernández-Alcaraz, M., Valencia-García, R., Meehan, K., & Lunney, T. (2018). Tourism recommender systems: An overview of recommendation approaches. International Journal of Computers and Applications,180(20), 9–13. Lavanya, D., & Rani, K. U. (2012). Ensemble decision making system for breast cancer data. International Journal of Computer Applications,51(17), 19–23. Le, Q. T., & Pishva, D. (2016). An innovative tour recommendation system for tourists in Japan. In 2016 18th international conference on advanced communication technology (ICACT) (pp. 717–729). IEEE. McF (1973). Conditional logit analysis of qualitative choice behavior. In Frontiers in econometrics. Academic Press. Mottini, A., Lhéritier, A., Acuna-Agost, R., & Zuluaga, M. A. (2018). Understanding customer choices to improve recommendations in the air travel industry (pp. 28–32). RecTour@ RecSys. Peska, L., & Vojtas, P. (2017). Using implicit preference relations to improve recommender systems. Journal on Data Semantics,6(1), 15–30. Pestana, M. H., Parreira, A., & Moutinho, L. (2020). Motivations, emotions and satisfaction: The keys to a tourism destination choice. Journal of Destination Marketing & Management,16(100), 332. Polikar, R. (2006). Ensemble based systems in decision making. IEEE Circuits and Systems Magazine,6(3), 21–45. Polydoropoulou, A., & Lambrou, M. A. (2012). Development of an e-learning recommender system using discrete choice models and bayesian theory: A pilot case in the shipping industry. In Security Enhanced Applications for Information Systems. IntechOpen. Quinlan, J. R. (1986). Induction of decision trees. Machine Learning,1(1), 81–106. Resnick, P., Iacovou, N., Suchak, M., Bergstrom, P., & Riedl, J. (1994). Grouplens: An open architecture for collaborative filtering of netnews. In Proceedings of the 1994 ACM conference on computer supported cooperative work (pp. 175–186). Saavedra, P., Barreiro, P., Duran, R., Crujeiras, R., Loureiro, M., & Vila, E. S. (2016). Choice-based recommender systems (pp. 38–46). RecTour@ RecSys. Salton, G., & McGill, M. J. (1986). Introduction to modern information retrieval. McGraw-Hill, Inc. Santos, F., Almeida, A., Martins, C., Gonçalves, R., & Martins, J. (2019). Using poi functionality and accessibility levels for delivering personalized tourism recommendations. Computers, Environment and Urban Systems,77(101), 173. Sen, A. (1990). Rational behaviour. In Utility and probability (pp. 198–216). Springer. Train, K. E. (2009). Discrete choice methods with simulation. Cambridge University Press. Utku, A., Karacan, H. U., Yildiz, O., & Akcayol, M. A. (2015). Implementation of a new recommendation system based on decision tree using implicit relevance feedback. JSW,10(12), 1367–1374. Wan, L., Hong, Y., Huang, Z., Peng, X., & Li, R. (2018). A hybrid ensemble learning method for tourist route recommendations based on geo-tagged social networks. International Journal of Geographical Information Science,32(11), 2225–2246. AUTHOR BIOGRAPHIES Ameed Almomani was born in Irbid, Jordan in 1980. He graduated from Yarmouk University (YU) in Jordan in 2002 in the field of computer science. After that, he studied a master's degree in computer information systems in Jordan as well, and graduated in 2006. He received a PhD in Artificial Intelligence from CiTIUS, one of the research centers of Santiago de Compostela University (USC). He worked in the Ministry of Education in Jordan and then as a lecturer at several universities in the Kingdom of Saudi Arabia. Nowadays, he is working with Artgro company as a Datascientist in the USA. His main research interests are in Artificial Intelligence, Recommender Systems, Machine learning, Choice modeling, and decision making. Paula Saavedra (Santiago de Compostela, 1984) is an Assistant Professor at the Department of Statistics, Mathematical Analysis and Optimization (USC). She obtained her PhD in Statistics and Operations Research in March 2015, with a thesis entitled “Nonparametric data‐driven methods for set estimation”(supervised by Prof. Wenceslao González‐Manteiga [USC] and Alberto Rodríguez‐Casal [USC]). Her main research is focused on nonparametric statistics methods for set estimation. Pablo Barreiro received his Master Degree in Information Technologies in 2014. He worked as a research assistant at the University of Santiago de Compostela (Spain), in the field of choice‐based recommendation algorithms. Nowadays, he left the academic world to join the software development industry, where he is working as a Senior Software Engineer at Auctane. Roi Durán received his Phd in Economics, with emphasis in environmental economics, University of Vigo. Currently, he works as a teacher of Secondary Education. Rosa Crujeiras received her Ph.D. degree in Mathematics in 2007. She has been a postdoc researcher at Universitè Catholique de Louvain (Belgium) and currently she is an associate professor at Universidade de Santiago de Compostela (Spain). She is the Scientific Director of the ALMOMANI ET AL.17 of 18 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Galician Centre for Mathematical Research and Technology (CITMAga). Her main research interest are focused on nonparametric statistical methods for dependent and complex data. María Loureiro is a Professor of Economics at University of Santiago de Compostela, Spain; PhD from Washington State University, USA. Researcher included in the Highly Cited Edition 2012, and on the top 2% most influential researchers according to Ioannidis JPA, Boyack KW, Baas J (2021) “Updated science‐wide author databases of standardized citation indicators.”PLoS Biol 18(10): e3000918. https://doi.org/10. 1371/journal.pbio.3000918. See recent data at https://elsevier.digitalcommonsdata.com/datasets/btchxktzyw/3. Eduardo Sánchez is an Associate Professor at the University of Santiago de Compostela. He obtained a Ms in Neuroscience at the International University of Andalucía, and a Ms in Computer Science and Software Engineering at the University of Southern California. He received his PhD degree in Physics in 2001. His research is focused in the fields of decision‐making and recommender systems, developing predictive models in the domains of tourism and culture, and computational neuroscience, working with models of the visual system. How to cite this article: Almomani, A., Saavedra, P., Barreiro, P., Durán, R., Crujeiras, R., Loureiro, M., & Sánchez, E. (2023). Application of choice models in tourism recommender systems. Expert Systems,40(3), e13177. https://doi.org/10.1111/exsy.13177 18 of 18 ALMOMANI ET AL. 14680394, 2023, 3, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13177 by Universidade de Santiago de Compostela, Wiley Online Library on [15/06/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License