scieee AI-readable full text Open interactive document viewer

Fuzzy metatopics predicting prices of Airbnb accommodations

Sánchez Franco, Manuel Jesús; Troyano Jiménez, José Antonio; Alonso-Dos-Santos, Manuel

Abstract

The purpose of this study is to guide pricing policies of Airbnb accommodation rentals to reduce inefficient pricing strategies through a novel application of topic modelling and a fuzzy clustering. In particular, the method proposes the application of Structural Topic Modelling, which explains a set of observations from latent topics. The associations between topics by Fuzzy C-Means Clustering are analysed to obtain new, more compact representations of topics (i.e., metatopics). This research identifies 15-metatopics related to Airbnb accommodations based on location and connectivity, enjoyment of domestic and everyday services, and the possibility of more authentic local experiences, among others. The influence of key metatopics on the price of Airbnb accommodations is determined by applying Extreme Gradient Boosting (an efficient and scalable implementation of gradient boosting framework) and Shapley Additive Explanations values. To sum up, our research provides an explicit contribution of user-generated content to promote the development of mutually beneficial relationships between guests and hosts, and detects future lines of research and practical and conceptual implications of the findings.

Full text

11/1/24, 15:06 Fuzzy metatopics predicting prices of Airbnb accommodations - IOS Press https://content.iospress.com/articles/journal-of-intelligent-and-fuzzy-systems/ifs189193 1/2 Fuzzy metatopics predicting prices of Airbnb accommodations Issue title: Fuzzy systems and applications in innovation and sustainability Guest editors: Ernesto Leon-Castro, Fabio Blanco-Mesa, Victor Alfaro-Garcia, Anna M. Gil-Lafuente and Jose M. Merigo Article type: Research Article Authors: Sánchez-Franco, Manuel J. (https://content.iospress.com:443/search?q=author%3A%28%22Sánchez-Franco, Manuel J.%22%29) | Troyano, José A. (https://content.iospress.com:443/search?q=author%3A%28%22Troyano, José A.%22%29) | Alonso-Dos-Santos, Manuel (https://content.iospress.com:443/search?q=author%3A%28%22AlonsoDos-Santos, Manuel%22%29) Af!liations: [a] Business Administration and Marketing, University of Sevilla, Spain | [b] Computer Languages and Systems, University of Sevilla, Spain | [c] Departamento de Administración, Universidad Católica de la Santísima Concepción, Concepción, Chile | [d] Department of Marketing and Market Research, University of Granada, Spain. Correspondence: [*] Corresponding author. Manuel J. Sánchez-Franco, Business Administration and Marketing, University of Sevilla, Spain. E-mail: [email protected] (mailto:[email protected]). Abstract: The purpose of this study is to guide pricing policies of Airbnb accommodation rentals to reduce inef!cient pricing strategies through a novel application of topic modelling and a fuzzy clustering. In particular, the method proposes the application of Structural Topic Modelling, which explains a set of observations from latent topics. The associations between topics by Fuzzy C-Means Clustering are analysed to obtain new, more compact representations of topics (i.e., metatopics). This research identi!es 15metatopics related to Airbnb accommodations based on location and connectivity, enjoyment of domestic and everyday services, and the possibility of more authentic local experiences, among others. The in"uence of key metatopics on the price of Airbnb accommodations is determined by applying Extreme Gradient Boosting (an ef!cient and scalable implementation of gradient boosting framework) and Shapley Additive Explanations values. To sum up, our research provides an explicit contribution of user-generated content to promote the development of mutually bene!cial relationships between guests and hosts, and detects future lines of research and practical and conceptual implications of the !ndings. Keywords: Sharing economy, Airbnb, price, fuzzy c-means clustering, structural topic model DOI: 10.3233/JIFS-189193 Journal: Journal of Intelligent & Fuzzy Systems (https://content.iospress.com:443/journals/journal-of-intelligentand-fuzzy-systems), vol. 40, no. 2, pp. 1879-1891, 2021 Published: 02 February 2021 Price: EUR 27,50 a; * b c; d FUZZY METATOPICS PREDICTING prices OF AIRBNB ACCOMODATIONS. Manuel J. SÁNCHEZ-FRANCO 000-0002-8042-3550 1 , José A. TROYANO-JIMÉNEZ 0000-0002-93173626 2 , Manuel ALONSO DOS SANTOS 0000-0001-9681-7231 3 1 Administración de Empresas y Marketing, Universidad de Sevilla, Sevilla, Spain 2 Lenguajes y Sistemas Informáticos, Universidad de Sevilla, Sevilla, Spain 3 Administración, Universidad de Católica de la Santísima Concepción, Concepción, Chile Corresponding author: Manuel J. Sánchez-Franco, [email protected] Received D M A; accepted date D M A Abstract. This research analyses the metatopics related to Airbnb accommodations based on location and connectivity, enjoyment of domestic and everyday services, and the possibility of more authentic local experiences, among others. Firstly, the method focuses essentially on the application of structural topic modelling which explains a set of observations from unobserved topics. Secondly, our research analyses the associations between topics by Fuzzy CMeans Clustering to obtain new, more compact representations of topics (i.e., metatopics) than the representations initially provided by single topics. Thirdly, our research determines the key metatopics that influence the price of Airbnb accommodations by applying Extreme gradient boosting (an efficient and scalable implementation of gradient boosting framework) and SHAP values. To sum up, our research (1) provides an explicit contribution of user-generated content to promote the development of mutually beneficial relationships between guests and hosts, (2) contributes to the literature on user behaviour in online home-sharing services and (3) detects future lines of research and managerial implications of service recommendation systems. Keywords: Airbnb; sharing economy; price; structural topic models; fuzzy c-means clustering; XGBoost; SHAP values. Introduction Dynamic pricing, infomediation or recommendation systems based on connecting supply and demand through peer-to-peer systems (1) modify the way in which individuals contract out hospitality services (cf. Raguseo et al., 2017; Sánchez-Franco et al., 2018; Sparks et al., 2016, among others), and (2) increase the relevance of strategies of value for money and pricing (Gibbs et al., 2018, p.46). Precisely, individuals employ recommendation systems (1) to reduce the risk of expectation failure by consulting freely sharing accommodations’ scores and travellers’ comments, and also (2) to generate by themselves a new content on their subjectively experienced encounters (hereinafter, UGC; e.g. Boyd and Ellison, 2008; Chevalier and Mayzlin, 2006; Mao and Liu, 2017; Moe and Schweidel, 2012; Mudambi and Schuff, 2010, among others). Although UGC is poorly structured and overall focused on a single entity or aspect of hospitality services, or is multi-lingual, it is mainly based on the authenticity, and helps to understand guests’ true preferences. UGC tends to be even more empathetic and trustworthy than other social communications or unidimensional metrics (Bickart and Schindler, 2001; Gretzel and Yoo, 2008; Sparks et al., 2016); for instance, “simple scores oversimplify quality measures by assuming that quality is a unidimensional measure” (Lawani et al., 2019, p.22; see also Archak et al., 2011); moreover, scores tend to be extremely high and promote loss of informative value, and do not appear to have a significant effect on listings' prices (Ert et al., 2015). In this regard, the aim of this paper is (1) to evidence the authentic interests extracted from customer reviews, (2) to guide pricing policies of accommodation rentals to reduce inefficient pricing strategies, and (3) to foster higher value for money. Our research specifically focuses on Airbnb that differs from traditional hotels in terms of booking systems, facilities, software platforms and design, and service for guests. Firstly, since launching in 2008, Airbnb has become one of the largest single tourism distribution platforms for short-term accommodation rentals (Gibbs et al., 2018). Individuals grant each other temporary access to underutilised physical assets, usually for money (Frenken and Schor, 2017; cf. also Sundararajan, 2016). In this regard, Airbnb is classified under the “Sharing Economy” (or collaborative consumption; cf. Frenkel et al., 2015). Secondly, customers’ interests extracted from UGC reveal, for instance, their preferences between different Airbnb lodgings. In this regard, our research focuses on analysing subjective experiences related to measurable, utility, emotional and social features from guests’ perspectives; e.g., economic benefits / price value (e.g., Mao and Lyu, 2017; Tussyadiah, 2015; Yang 2 and Ahn, 2016), authenticity (e.g., Guttentag et al., 2018; Poon and Huang, 2017), home benefits (e.g., Guttentag et al., 2018; Johnson and Neuhofer, 2017), social interactions and benefits enjoyed from using Airbnb (e.g., Moody et al., 2017; Tussyadiah and Pesonen, 2016), or distance from the touristy centre (cf. also Guttentag, 2016; Satama, 2014; So et al., 2018, among other authors). Although previous studies examine the effects of number of reviews, ratings, and host photos on the prices of Airbnb lodgings, scarce research focuses on guests’ expectations, predictions, goals, and desires from linguistic attributes of online textual reviews generated by customers to determine holistically the price of sharing economy (Zhao et al., 2019). Accordingly, our research applies a product, feature-oriented approach to identify relevant topics (a distribution of terms over a fixed vocabulary, Blei, 2012; see also Blei et al., 2003) around (1) the core and basic sharing services (here, Airbnb lodgings) as well as (2) surrounding features (e.g. neighbourhood amenities or tourist hotspots, among others) that significantly affect the prices of accommodation listings and generate higher revenues for hosts. Our method lies in improving the performance of guests’ reviews on prediction tasks by (1) identifying users’ experience-related frequent terms and relevant topics, (2) examining the underlying semantic structure and reducing the number of topics into meaningful fuzzy clusters (or metatopics) that make them easier to interpret, (3) determining the significant metatopics among a large quantity of text data that influence the price of Airbnb accommodations, and consequently, (4) providing an explicit contribution of UGC to promote the development of mutually beneficial relationships. Our paper is structured as follows. The second section presents the theoretical background derived from prior literature. The third section elaborates the method, and the results of the study that provide a new framework to understand comprehensively the influential drivers of Airbnb pricing. The last section concludes with implications and limitations. 1. Theoretical framework Peer-to-peer economy or collaborative consumption are based on network effects that enable hosts to offer, for instance, their unoccupied houses or rooms (to other individuals) for short-term rental, near-zero marginal costs, and reduced search costs (cf. Zervas et al., 2017; Rifkin, 2014). In particular, Airbnb is primarily a low-cost option where travellers (or here, guests) easily find entire apartments or rooms at a more competitive price than hotels coordinated through community-based services, and “(…) can easily price-shop multiple accommodation options with the click of the mouse” (Gibbs et al., p.46). Moreover, the demand for Airbnb rentals is significantly elastic because of the hosts’ fixed costs of rent and utilities, together with minimal labour costs and probably untaxed extra income (cf. Oskam and Boswijk, 2016). Although guests or hosts do not have significant market power and thus Airbnb accommodation prices become one of the key determinants that affects guests’ selection of them and warrants appropriately their revenue (Guttentag and Smith, 2017; Lampinen and Cheshire, 2016; Wang and Nicolau, 2017), few studies have been conducted to identify what features influence pricing policies of Airbnb accommodations. Firstly, previous research has detected inefficient pricing by Airbnb hosts due to (1) uniqueness of the rental services offered on Airbnb (Gibbs et al., 2018) and (2) emotional drivers applied by non-professional hosts (cf. Ikkala and Lampinen, 2014; see also Hill, 2015, Li et al., 2016, among others). To solve partially this research gap, our study extracts rational attributes and affective or social topics from customer reviews (guests’ perspective) and analyses the full dynamics that most often influence prices of Airbnb accommodations in a comprehensive pricing analysis. Indeed, to ignore topics from guests’ narratives (defined as a mixture over terms where each term has a probability of belonging to a topic k; cf. Blei, 2012 or Blei et al., 2003, among others) can yield inaccurate estimates of the prices, and the size of the pricing errors can affect the true conclusions about the research implication on guests’ welfare. Secondly, a pricing policy for hosts should entail the assessment of which utility-bearing attributes or quality signs, and to what extent they matter to guests, to gain value for their expenditure. Hosts should know the preferences that drive guests’ decisions. Accordingly, our research proposes that the price of an Airbnb accommodation should be a Hedonic Pricing Analysis-based function related to the presence or absence of relevant and essential ‘subjective dimensions or features’ (hereinafter, HPA; see Lancaster’s theory of consumer demand; cf. also Rosen, 1974). On the one hand, in the hospitality domain HPA has been widely used to determine the price of tourism packages (Aguiló et al., 2005; Clewer et al., 1992; Mangion et al., 2005; Taylor, 1995), urban hotels (Chen and Rothcschild, 2010; Thrane, 2007; Zhang et al., 2011), hotel rooms (cf. Espinet et al., 2003; Rigall-i-Torrent et al., 2011), private rental properties in holiday domains (cf. Hamilton, 2007; Portolan, 2013), or Bed & Breakfasts (cf. Monty and Skidmore, 2003). Also, Ert et al. (2016) and Teubner et al. (2017) focus on the photographs that owners upload to the platform, their star rating or owner response. HPA is also applied to 3 estimate the value of different attributes in price composition on the Airbnb services by studies such as Chen and Xie (2017), Gibbs et al. (2018), or Wang and Nicolau (2017), among others. Thirdly, assuming that “an Airbnb accommodation listing is a bundle of elements that influence the quality of the overall product and provide consumers with value and satisfaction” (Gibbs et al., p.48), previous research provides valuable discussions about the presence or absence of key subjective dimensions or features and their contributions on model output. In particular 1 : ˗ Site-specific features measured as distance from the touristy hotspots (or also a major transportation hub), or hospitality features based on home-like lodging conditions (e.g. household amenities and basic functionalities such as homely feel, real beds, wireless Internet, large space, and free parking, among others) are traditional factors that determine a customer's final choice (cf. Guttentag, 2015, 2016; Johnson and Neuhofer, 2017). ˗ Airbnb guests also focus on the emergence of a society that desires value for money (based on more transparent pricing; cf. Guttentag, 2016; Mao and Lyu, 2017; Satama, 2014; Tussyadiah and Pesonen, 2016; Yang and Ahn, 2016) or postmodern experiences described as authentic staying at an Airbnb lodging (cf. Guttentag et al., 2018; Liang, 2015; Mody et al., 2017; Poon and Huang, 2017), novelty (Guttentag, 2016; Johnson and Neuhofer, 2017; Mao and Lyu, 2017), or interaction as part of a social benefit from using Airbnb (Tussyadiah and Pesonen, 2016). Airbnb is indeed focused on postmodern tourists, in contrast to modern tourists, i.e. travellers who enjoy multiple experiences embracing different, sometimes contrasting life values, or interaction with the host and local people (Guttentag, 2016; Johnson and Neuhofer, 2017; Mody et al., 2017; Poon and Huang, 2017; Tussyadiah and Pesonen, 2016). Although guests' motivations for selecting Airbnb accommodations have been researched by a handful of scholars, “the research to date related to ‘pricing and Airbnb’ does little to explain the variables that make up the price of a listing” (Gibbs et al., 2018, p.47). It mainly focuses on examining a few influential drivers in isolation, without providing a broad perspective on the issue (So et al., 2018); moreover, “this body of research also suffers from numerous limitations, (…) and the studies reach somewhat incongruent conclusions” (Guttentag et al, 2018, p.344). 2. Materials and methods Airbnb accommodation rental offers are analysed using an HPA in the five most populated New York urban neighbourhoods. New York city registers 65.2 million visitors (NYC & Company, 2019), and is selected as a case study because it generates rich feelings that can affect guests' perceptions. The dataset is obtained from the Inside Airbnb website (http://insideairbnb.com/ -a non-commercial, open source data tool on Airbnb). 2.1. Data collection Initially the dataset contains: ˗ 1,106,639 reviews located exclusively in New York to avoid any risk of heterogeneity induced by ‘regional’ effect that prevents relevant possible comparisons. ˗ 49,748 listings - distributed in the following neighbourhoods: Bronx (1,001), Brooklyn (20,312), Manhattan (22,559), Queens (5,525), and Staten Island (351). ˗ Airbnb listings managed by 37,689 hosts in 193 census tracts (zip codes). Although the use of five neighbourhoods within one city provides a dataset which has similarities in culture, Airbnb rules, and travel seasonality (Gibbs et al., 2018), in order to standardise the comparisons and to increase the comparability among listings, our research includes accommodations for only (1) less than six guests, and (2) entire apartments during three consecutive years, from 2016 to 2018. Furthermore, our research analyses a single language, English, to keep the language variable consistent across texts, and it applies the textcat package based on the R 3.6.1 statistical tool to recognise English in the reviews (cf. Hornik et al., 2013), and Google's Compact Language Detector 2. 1 Cf. also So et al. (2018), Chan and Wong (2006), Gutt and Herrmann (2015), Li et al. (2016) Li and Tabari (2019), Panda et al. (2015), Stors and Kagermeier (2015), Tussyadiah (2016), and Wang and Nicolau (2017). 4 2.2. Data cleansing process Our research carries out a cautious data cleaning process (to increase the quality of the metatopics) based on transforming free-form text into a structured form. It applies the following stages: (1) it discards punctuation, capitalisation, digits, and extra whitespace, (2) recognises common abbreviations and acronyms (e.g., xmas or Christmas), (3) removes a list of stop words to filter out overly common terms without specific relevance for the research problem, and (4) tokenises and lemmatises the terms. To avoid tallying one term in various grammar contexts, our research only retains the stem of a term. It also omits terms shorter than a minimum of three characters. Finally, our dataset contains 40,572 reviews (and 9,710 listings). The numerical ratings average 94.47 (with sd = 4.51, and a minimum, median, and maximum of 20, 95, and 100, respectively). Price is taken at the accommodation level and does not include cleaning fees or additional charges for guests that are not included in the overall price. Removing outliers, the average price is $157.59 (sd = 58.34 with a minimum, median, and maximum of $10, $150, and $320, respectively). 2.3 Extracting terms Our initial corpus contains more than 292,253 lemmatised terms (nouns and verbs). Applying a topic modelling on all terms in a corpus is both computationally expensive and not very useful. The inclusion of redundant, irrelevant and noisy terms in topic building process could also cause poor predictive performance. In this vein, our research selects a subset of terms that minimises the redundancy and maximises their relevance. Online reviews overall tend to be brief with only a relatively small number of major topics standing out, and usually contain no terms that occur more than once per document. Our research then applies Bi-Normal Separation metric (hereinafter, BNS; see Forman, 2008) that does not rely on term frequency within the document, enhances vector representations of short documents, and outperforms tf-idf scaling algorithms for prediction tasks. This process yields a dictionary of 912 terms. 2.4. Data mining Although estimating the BNS values is useful in knowledge extraction, to identify the rational and experiential topics in the corpus and their associations is a more powerful approach for understanding the true-context of the hospitality-based opinions. Next, our research estimates the relationships between terms and documents through a text-mining algorithm to discover hidden semantic structures in the corpus (cf. Gutiérrez et al., 2017). In particular, our research analyses natural and non-structured reviews by machine-learning algorithms based on text summarisation and the application of structural topic modelling (hereinafter, STM; cf. Roberts et al., 2013). 2.4.1. Structural topic model: Model specification and selection Topic modelling is nowadays a computer-assisted technique that can help scholars and managers to address the costs and time associated with the growing amount of data and uncovers patterns of term co-occurrence across the corpus. “Our goal is here to describe the data with fewer dimensions (topics) than are actually present, but with enough dimensions so that as little relevant information as possible is lost” (Jacobi et al., 2015, p.5). Our approach focuses on STM -a generative model of term counts-, and its implementation in the STM 1.3.3 R package (Roberts et al., 2018). STM, as an unsupervised method, allows the researchers to discover topics (inferred here from the guests’ narratives) that can be correlated, and estimates their relationships to document metadata (explanatory covariates defined as information about each document, e.g. price). On the one hand, each review is a mixture of topics, and each topic corresponds to a different categorical distribution of terms (cf. also Latent Dirichlet Allocation, Blei et al. 2003). On the other hand, STM can estimate the impacts of metadata on topic prevalence to take the context into account for a better understanding of the ‘semantically interpretable themes’ without forcing the metadata to be influential on the extracted topics. Although “for large corpora (…) previous research finds that between 60 and 100 topics are best” (cf. Lindstedt, 2019, p.311), our research prioritises managerial implications, and consequently, proposes extracting a smaller number of topics. In this regard, hosts prefer prediction models that not only provide technical insights but are also accessible and offer easy to understand managerial implications. Assuming, therefore, that there is no single correct path to select an optimal number of topics, our research here (1) estimates different STM models for 30 and 50 topics; (2) proposes an initialisation based on the method of moments, “which is deterministic and globally consistent under reasonable conditions” (Roberts et al., 2018, p.11; cf. also Roberts et al. 2016); (3) discards the 5 models that have the lowest value for the bound (Roberts et al., 2014), and (4) assesses the trade-off between semantic coherence and exclusivity (i.e. internal consistency -cohesivenessand differentiation or diversity criteria, see Blei et al., 2003; Gerring, 2001, Mimno et al., 2011; Roberts et al., 2014). Accordingly: ˗ This process results in selecting a subset of models with average scores towards the upper right side of the plot prioritising partially diversity criteria (here, 39-45 topics). Fig. 1 shows that overall the semantic values decrease (since more topics implies fewer co-occurrences of terms), and contrariwise, the exclusivity values as a whole increase. ˗ Our research prioritises that extracted topics are capturing different conceptual aspects of Airbnb experiences (cf. discriminant validity). It fits into “recent research which has begun to show diversity in motivations for participating in the sharing economy” (cf. Lutz and Newlands, 2018, p.188). Therefore, models with less than 38 topics, which score highly on semantic coherence but less on exclusivity, or scored equally, are here likely not to be a suitable solution to reveal functional, affective and social hospitality-experiences. ˗ The highest exclusivity value within this grouping (39-40 topics) is the case where 44 topics are revealed. Our research eventually selects a 44-topic model to start to estimate the “true” number of topics and to ensure both interpretable and useful results (see Fig. 2). It is thus an acceptable (not optimal) number of topics for uncovering useful features about the natural content of the guests’ reviews in relation to the Airbnb experiences. ˗ In order to select a definitive model based on its predictive validity, our research assesses (a) the candidate models’ outputs with highest capacity to predict price accommodations by r2 metric from applying an extreme gradient boosting (hereinafter, XGBoost), and (b) the candidate models that are based on a manageable number of metatopics. Figure 1: Semantic Coherence and Exclusivity values for 200 runs of STM. 6 Figure 2: Example of graphic display of semantic coherence and exclusivity values for 44 topics. To identify intuitive meanings of topics, they are also conceptualised with a list of the most representative customer reviews that are most strongly connected to each topic. Although the STM proposal extracts the terms that have the highest probability of occurring conditionally on the topic within a model, selected terms may not be semantically interesting (Kuhn, 2018). Bischof and Airoldi (2012) propose using the FREX statistic that combines term frequency and exclusivity to topics. FREX is defined as the ratio of term frequency conditional on a topic to term-topic exclusivity (cf. Roberts et al., 2013; Roberts et al. 2014). ω weight balances the influence of frequency and exclusivity, and it is here set to 0.5. Fig. 3 provides a summary of 44-topics as an illustrative proposal, their mean prevalence and the most FREX terms. The topics 3 and 23 related to home benefits (around 4.5% of the documents) and ‘connectivity’ (4.3%) show the most estimated topic proportions. To validate how valuable as subject matter scholars and managers find extracted topics (Chang et al., 2009), it is necessary initially to compare them with other studies (see Epigraph 1: Theoretical framework). Our research also focuses on the model’s semantic validity, i.e., “the extent to which each category or document has a coherent meaning and the extent to which the categories are related to one another in a meaningful way” (Quinn et al., 2010, p.216), and on predictive validity (i.e., the extent to which the measure corresponds correctly to external events): ˗ Following Quim et al. (210), “the coherent meaning of the metatopics we find is further evidence of the semantic validity of the topic model” (p.218). One way to define topics is here to categorise them by clusters revealing their organisation to examine the semantic relationships within and across clusters of topics (hereinafter, metatopics) by applying a fuzzy clustering analysis and by using of the theta matrix based on the posterior probability of a topic given a document as clustering inputs. ˗ For predictive validity, our research analyses the metatopics’ influence on price of accommodations by applying XGBoost. 7 Figure 3: Example of graphical display of estimated topic proportions ( ω = 0.5). 2.4.2. Topic groupings: Fuzzy clustering analysis Overall cluster analysis provides a set of methods and tools to gain a deeper understanding of a system by discerning important grouping patterns. Clusters detection is indeed employed in order to reduce the number of topics into meaningful groupings of them (metatopics) that would be useful to interpret their influence on price of accommodations. Moreover, clustering is performed by pursuing different potential benefits, i.e., eliminating noisy features, avoiding overfitting, or summarising original features to make the human interpretation of results easier. In the latter case, minor loss in the predictive capacity of the new features is acceptable if our results improve interpretability of results. Research identifies two main groups of dimensionality reduction techniques; feature selection and feature extraction. On the one hand, feature selection tries to find the subset of original features that best captures the information contained in the dataset. On the other hand, feature extraction defines a new set of features calculated from the original features. Precisely, our research applies feature extraction technique based on Fuzzy C-Means Clustering (hereinafter, FCM as an extension of the hard K-means algorithm to the fuzzy framework). FCM obtains new, more compact representations of documents than those initially provided by extracted topics; in particular, FCM allows the researcher to deal with lexical ambiguity because a single term (or a topic) could belong to multiple semantic categories. Despite research regarding fuzzy clustering conducted in the past, scarce attention is paid to its applications in tourism (d’Urso et al., 2016). FCM is initially studied by Dunn (1974) and generalised by Bezdek in 1974 (Bezdek, 1981; Bezdek et al., 1984). FCM is inspired by the hard-clustering algorithm called K-means and is based on the concept of centroid. In Kmeans, centroids are calculated using the average of each of the instances belonging to the cluster; meanwhile, in FCM centroids are calculated with a weighted average (using membership degrees as coefficients) of a transformation of instances. Transformations consist in raising each instance using the parameter m as a fuzziness exponent or fuzzification degree; i.e., higher values of m cause lower degrees of membership and subsequently, higher fuzzy partitions. m is a real number greater than 1. In comparison to hard clustering or crisp clustering, each review is a mixture of topics and each topic is thus the member of distinct clusters with varying degrees of membership between 0 and 1. FCM designs the following iterative scheme for a given number of clusters c: Assign random membership degree for each instance and each cluster Repeat until converge (membership degrees stabilise) 8 Compute centroids for each cluster For each instance Compute new membership degrees based on the new centroids As Fig. 4 graphically displays, our research groups the features (i.e., dataset columns) instead of instances (i.e., dataset rows). Our interest lies in distributing the information given by each feature in different clusters and keeping the centroids as representatives of extracted clusters acting as new features. Figure 4: Customised process of fuzzy clustering of topics. To sum up, FCM is more capable of analysing the uncertainty and vagueness that characterise the experiences and overlapping perceptions of postmodern consumers (d’Urso et al., 2016), and subsequently, to enhance the predictive capacity of subjective features-based communities on the prices of accommodation listings. 2.5. Predictive analysis: Extreme gradient boosting Although generalised regression models such as OLS regression or quantile regression are the most commonly adopted methods to analyse features affecting accommodation prices, our research here applies XGBoost. It is initially proposed by Friedman (2001), and it starts as a research project by Tianqi Chen (Chen and Guestrin, 2016). On the one hand, boosting is an ensemble technique that relies on the idea that a series of weak estimators (classifiers or regressors) can behave like a robust estimator. One of the bases to converge on a robust solution is the application of an iterative and adaptive scheme, in which each new weak classifier mainly concentrates on instances of data that have been misclassified by previous classifiers; i.e., every new tries to correct the errors of previously tries. On the other hand, XGBoost is an efficient and scalable implementation of gradient boosting framework, or GBM (Climent et al., 2019; Chen and Guestrin, 2016). GBM is the most popular classifier, and it has the particularity of interpreting the boosting process in terms of the optimisation of a cost function, which allows the use of an adaptation of the gradient descent algorithm for guiding the training. The base estimators are usually shallow decision trees and, unlike other combination schemes such as random forests, the models obtained with gradient boosting are usually free of overfitting. XGBoost includes a series of optimisations that make the training much faster than other implementations, along with regularisation techniques that help reduce overfitting. 3. Experimental results Our main goal is to apply FCM (1) to obtain a new dataset with fewer features (metatopics) that gathers most of the information from the original dataset, and (2) to predict their influence on price of accommodations. The quality of the reductions is based on the performance of an XGBoost regressor trained with the new features, where All models are fitted in R (R-3.6.1); in particular, for XGBoost applying (1) xgboost package version 0.90.0.2 (Chen et al., 2019) and (2) caret package version 6.0–84 (Kuhn, 2019). 3.1. External and internal evaluation Our research uses the coefficients root mean squared error (RMSE) and r-squared (r2) scores to validate the models and evaluate their qualities. r2, in particular, is one of the most common metrics to evaluate regressors. To ensure the effectiveness of the training process and to find the best fitted model based on r2 metric, our research splits the dataset into two subsets: Dataset Centroids DatasetTraspose Original features Fuzzy clustering Traspose New features 15 Cited references ˗ Aguiló, E., Alegre, J. & Sard, M. 2003. Examining the market structure of the German and UK tour operating industries through an analysis of package holiday prices. Tourism Economics, 9 (3), 255-278 ˗ Archak, N., Ghose, A. & Ipeirotis, P.G. 2011. Deriving the pricing power of product features by mining consumer reviews. Management Science, 57, 1485-1509 ˗ Bezdek, J. C., Ehrlich, R. & Full, W. 1984. FCM: The fuzzy c-means clustering algorithm. Computers & Geosciences, 10 (2-3), 191-203. ˗ Bezdek, J.C. 1974. Cluster validity with fuzzy sets. Journal of Cybernetics, 3 (3), 58–73. ˗ Bezdek, J.C. 1981. Pattern recognition with fuzzy objective function algorithms. Kluwer Academic Publishers, Norwell. ˗ Bickart, B. & Schindler, R. M. (2001). Internet forums as influential sources of con-sumer information. Journal of Interactive Marketing, 15, 31–40. ˗ Bischof, J. & Airoldi. E. 2012. Summarizing topical content with word frequency and exclusivity. In John Langford and Joelle Pineau (eds.), Proceedings of the 29th International Conference on Machine Learning (ICML-12). New York, NY: Omnipress, 201–208. ˗ Blei D.M. (2012) Probabilistic topic models. Surveying a suite of algorithms that offer a solution to managing large document archives. Communications of the ACM, 55 (4), 77–84 ˗ Blei, D.M., Ng, A.Y. &. Jordan, M.I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3: 9931022. ˗ Boyd, D.M. & Ellison, N.B. 2008. Social Network Sites: Definition, history, and scholarship. Journal of ComputerMediated Communication, 13, 210-230. ˗ Chan, E. S. & Wong, S. C. (2006). Hotel selection: When price is not the issue. Journal of Vacation Marketing, 12 (2), 142–159. ˗ Chang, J., Boyd-Graber, J., Wang, Ch., Gerrish, S. & Blei, D.M. 2009. Reading tea leaves: How humans interpret topic models. Advances in Neural Information Processing Systems, 288–296. ˗ Chen, C. & Rothschild, R. 2010. An application of hedonic pricing analysis to the case of hotel rooms in Taipei. Tourism Economics, 16 (3), 685–694. ˗ Chen, T. & Guestrin, C. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). ˗ Chen, T., He, T., Benesty, M., Khotilovich, V., & Tang, Y. 2019. XGBoost: eXtreme Gradient Boosting (R package version 0.90.0.2). Available at: https://https://cran.r-project.org/web/packages/xgboost/xgboost.pdf ˗ Chen, Y. & Xie, K. 2017. Consumer valuation of Airbnb listings: a hedonic pricing approach», International Journal of Contemporary Hospitality Management, 29 (9), 24052424. ˗ Chevalier, J. & Mayzlin, D. 2006. The effect of word of mouth on sales: Online book reviews. Journal of Marketing Research, 48 (August), 345-354. ˗ Clewer, A., Pack, A. & Sinclair, M.T. 1992. Price competitiveness and inclusive tour holidays in European cities. In Johnson, P., and Thomas, B., eds, Choice and Demand in Tourism, Mansell, London. ˗ Climent, F., Momparler, A. & Carmona, P. 2019. Anticipating bank distress in the Eurozone: An Extreme Gradient Boosting approach. Journal of Business Research, 101, 885-896, ˗ d`Urso P., Disegna, M., Massari, R. & Osti, L. 2016. Fuzzy segmentation of postmodern tourists. Tourism Management, 55, 297-308 ˗ Dunn, J.C. 1974. A Fuzzy Relative of the isodata process and its use in detecting compact well separated clusters. Journal of Cybernetics, 3 (3), 32-57. ˗ Ert, E., Fleischer, A. & Magen, N. 2016. Trust and reputation in the sharing economy: The role of personal photos on Airbnb. Tourism Management, 55 (1), 62–73. ˗ Espinet, J.M., Saez, M., Coenders, G. & Fluvia, M. 2003. Effect on prices of the attributes of holiday hotels: a hedonic prices approach. Tourism Economics, 9, 165–177. ˗ Forman, G. 2008. BNS feature scaling: an improved representation over TF-IDF for SVM text classification. In Proceedings of the 17th ACM Conference on Information and Knowledge Management, 263–270. ˗ Frenken, K. & Schor, J. 2017. Putting the sharing economy into perspective. Environmental Innovation and Societal Transitions, 23, 3-10 ˗ Frenken, K., Meelen, T., Arets, M. & Glind, P.V. 2015. Smarter regulation for the sharing economy. The Guardian, 20 May 2015. Available at: www.theguardian.com/science/political-science/2015/may/20/smarter-regulation-forthe-sharing-economy (accessed 30 July 2019). ˗ Friedman, J. 2001. Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29 (5), 1189-1232. ˗ Gerring, J. 2001. Social science methodology: A unified framework. Cambridge: Cambridge University Press. ˗ Gibbs, Ch., Guttentag, D., Gretzel, U., Morton, J. & Goodwill, A. 2018. Pricing in the sharing economy: a hedonic pricing model applied to Airbnb listings. Journal of Travel & Tourism Marketing, 35 (1), 46-56. ˗ Gretzel, U. & Yoo, K.H. 2008. Use and impact of online travel reviews. In P. O’Connor, W. Höpken, and U. Gretzel (eds.), Information and communication technologies in tourism, 35-46. Vienna: Springer. ˗ Gutierrez, J., García-Palomares, J.C., Romanillos, G. & Henar Salas-Olmedo, M. 2017. The eruption of Airbnb in tourist cities: Comparing spatial patterns of hotels and peer-to-peer accommodation in Barcelona. Tourism Management, 62, 278-291. 16 ˗ Gutt, D. & Herrmann, P. (2015). Sharing means caring? Hosts’ price reaction to rating visibility. In ECIS 2015 Research-in-Progress Papers (Paper 54). Retrieved from http://aisel.aisnet.org/ecis2015_rip/54/. ˗ Guttentag, D. 2015. Airbnb: Disruptive innovation and the rise of an informal tourism accommodation sector. Current Issues in Tourism, 18 (12), 1192-1217. ˗ Guttentag, D. 2016. Why tourists choose Airbnb: A motivation-based segmentation study underpinned by innovation concepts, (Unpublished Doctoral Dissertation), University of Waterloo, Canada. ˗ Guttentag, D. A. & Smith, S. L. 2017. Assessing Airbnb as a disruptive innovation relative to hotels: Substitution and comparative performance expectations. International Journal of Hospitality Management, 64, 1-10. ˗ Guttentag, D., Smith, S., Potwarka, L. & Havitz, M. 2018. Why Tourists Choose Airbnb: A motivation-based segmentation study. Journal of Travel Research, 57 (3), 342–359. ˗ Hamilton, J.M. 2007. Coastal landscape and the hedonic price of accommodation. Ecological Economics, 62 (3-4), 594602. ˗ Hill, D. 2015. How much is your spare room worth? IEEE Spectrum, 52 (9), 32-58. ˗ Hornik, K., Mair, P., Rauch, J., Geiger, W., Buchta, C. & Feinerer, I. 2013. The textcat package for n-gram based text categorization in r. Journal of Statistical Software, 52 (6), 1-17. ˗ Ikkala, T. & Lampinen, A. 2014. Defining the price of hospitality: Networked hospitality exchange via Airbnb. In The Proceedings of the companion publication of the 17th ACM conference on Computer supported cooperative work & social computing, February 15–19, Baltimore, Maryland, USA. ˗ Jacobi, C., van Atteveldt W. & Welbers, K. 2015. Quantitative analysis of large amounts of journalistic texts using topic modelling, Digital Journalism, 4 (1), ˗ Johnson, A.-G. & Neuhofer, B. 2017. Airbnb – an exploration of value co-creation experiences in Jamaica. International Journal of Contemporary Hospitality Management, 29 (9), 2361-2376. ˗ Kuhn, K.D. 2018. Using structural topic modelling to identify latent topics and trends in aviation incident reports, Transportation Research Part C: Emerging Technologies, 87 (February), 105-122, ˗ Kuhn, M. (2019). Caret: classification and regression training (R package version 6.0–84). Available at: https://cran.r-project.org/web/packages/caret/caret.pdf ˗ Lampinen, A. & Cheshire, C. 2016. Hosting via Airbnb: Motivations and financial assurances in monetized network hospitality. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, San Jose, CA, USA, 7–12 May, 1669–1680. ˗ Lancaster, K.J. 1971. Consumer Demand: A New Approach. New York: Columbia University Press. ˗ Lawani, A., Reed, M.R., Mark, T. & Zheng Y. 2019. Reviews and price on online platforms: Evidence from sentiment analysis of Airbnb reviews in Boston. Regional Science and Urban Economics, 75, March, 22-34 ˗ Li, J., Moreno, A., & Zhang, D. J. 2016. Pros vs Joes: Agent pricing behaviour in the sharing economy (SSRN Scholarly Paper No. ID 2708279). Rochester, NY: Social Science Research Network. ˗ Li, L. & Tabari, S. 2019. Impact of Airbnb on customers' behaviour in the UK hotel industry. Tourism Analysis, 24 (1), 13-26. ˗ Li, Y., Pan, Q., Yang, T. & Guo, L. 2016. Reasonable price recommendation on Airbnb using Multi-Scale clustering. In Proceedings of the 35th Chinese Control Conference, Chengdu, China, 27–29 July 2016, 7038–7041. ˗ Liang, L.J. 2015. Understanding repurchase intention of Airbnb consumers: perceived authenticity, EWoM and price sensitivity. (Unpublished Master's Thesis). University of Guelph, Canada ˗ Liang, L.J., Choi, Ch.C.H. & Joppe, M. 2018. Understanding repurchase intention of Airbnb consumers: perceived authenticity, electronic word-of-mouth, and price sensitivity. Journal of Travel & Tourism Marketing, 35 (1), 7389. ˗ Lindstedt, N.C. 2019. Structural Topic Modelling for social scientists: A brief case study with social movement studies literature, 2005–2017. Social Currents, 6 (4), 307–318. ˗ Lundberg S.M., Erion G.G. & Lee, S.I. 2018. Consistent individualised feature attribution for tree ensembles. [online] Available at: https://arxiv.org/pdf/1802.03888.pdf [Accessed 31 March 2018]. ˗ Lundberg, S.M. & Lee, S-I. 2917. A unified approach to interpreting model predictions. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 4765–4774. Curran Associates, Inc. ˗ Lundberg, S.M., Erion, G., Chen, H., DeGrave, A., Prutkin, J.M., Nair, B., Katz, R., Himmelfarb, J., Bansal, N. & Lee, S. 2019. Explainable AI for trees: From local explanations to global understanding. [online] Available at: https://arxiv.org/abs/1905.04610 [Accessed 31 March 2018]. ˗ Lutz, Ch. & Newlands, G. 2018. Consumer segmentation within the sharing economy: The case of Airbnb. Journal of Business Research, 88 (July), 187-196. ˗ Mangion, M.L., Durbarry, A & Sinclair, M. 2005. Tourism competitiveness: price and quality. Tourism Economics,1 (11),45-68. ˗ Mao, Z. & Lyu, J. 2017. Why travellers use Airbnb again? International Journal of Contemporary Hospitality Management, 29 (9), 2464-2482. ˗ Mimno, D., Wallach, H.M., Talley, E., Leenders, M. & McCallum, A. 2011. Optimizing semantic coherence in topic models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing EMNLP ’11, 262–72. Edinburgh, UK: Association for Computational Linguistics. ˗ Mody, M.A., Suess, C. & Lehto, X. (2017). The accommodation experiencescape: a comparative assessment of hotels and Airbnb. International Journal of Contemporary Hospitality Management, 29 (9), 2377-2404. 17 ˗ Moe, W. W. & Schweidel, D. A. (2012). Online product opinions: incidence, evaluation and evolution. Marketing Science, 31 (May-June), 372-386. ˗ Möhlmann, M. 2015. Collaborative consumption: determinants of satisfaction and the likelihood of using a sharing economy option again. Journal of Consumer Behaviour, 14(3), 193-207. ˗ Monty, B. & Skidmore, M. 2003. Hedonic pricing and willingness to pay for bed and breakfast amenities in Southeast Wisconsin. Journal of Travel Research, 42 (2), 195–199. ˗ Mudambi, S. & Schuff, D. 2010. What makes a helpful online review? A study of customer reviews on amazon.com. Management Information Systems Quarterly, 34 (1), 185–200. ˗ Oskam, J. & Boswijk, A. 2016. Airbnb: the future of networked hospitality businesses. Journal of Tourism Futures, 2 (1), 22-42. ˗ Panda, P., Verma, S. & Mehta, B. 2015. Emergence and acceptance of sharing economy in india: Understanding through the case of airbnb. International Journal of Online Marketing, 5, 3, 1-17 ˗ Poon, K. & Huang, W. 2017. Past experience, traveller personality and tripographics on intention to use Airbnb. International Journal of Contemporary Hospitality Management, 29 (9), 2425-2443. ˗ Portolan, A. 2013. The impact of the attributes of the private tourist accommodation facilities onto prices: A Hedonic price approach. European Journal of Tourism Research, 6 (1), 7482. ˗ Quinn, K.M., Monroe, B.L., Colaresi, M., Crespin, M.H. & Radev, D.R. 2010. How to analyse political attention with minimal assumptions and costs. American Journal of Political Science, 54 (1), 209-228. ˗ Raguseo, E., Neirotti, P. & Paolucci, E. 2017. How small hotels can drive value their way in infomediation. The case of ‘Italian hotels vs. OTAs and TripAdvisor’. Information & Management, 54 (6), 745-756. ˗ Rifkin, J. 2014. Uber and the zero marginal cost revolution. Huffpost 3 November, 2014. Available at: http://www.huffingtonpost.com/jeremy-rifkin/uber-german-court_b_5758422.html, (accessed 30 July 2019) ˗ Rigall-I-Torrent, R., Fluvià, M., Ballester, R., Saló, A., Ariza, E. & Espinet, J.M. 2011. The effects of beach characteristics and location with respect to hotel prices. Tourism Management, 32 (5), 1150-1158. ˗ Roberts M.B., Stewart B. & Tingley D 2016. Navigating the local modes of big data: The case of topic models. In Computational Social Science: Discovery and Prediction. Cambridge University Press, New York. ˗ Roberts, M.E., Stewart, B., and Tingley, D. 2018. stm: R Package for structural topic models. Journal of Statistical Software, forthcoming. ˗ Roberts, M.E., Stewart, B.M, Tingley, D., Lucas, C., Leder-Luis, J., Gadarian, S.K., Albertson, B. & Rand, D.G. 2014. Structural topic models for open ended survey responses. American Journal of Political Science, 58 (4), 10641082. ˗ Roberts, M.E., Stewart, B.M., Tingley, D. & Airoldi, E.M. 2013. The structural topic model and applied social science. In NIPS 2013 Workshop on Topic Models: Computation, Application, and Evaluation, pp. 2-5. ˗ Rosen, Sh. 1974. Hedonic prices and implicit markets: Product differentiation in pure competition. Journal of Political Economy, 82 (1), 34. ˗ Sánchez-Franco, M. J., Roldán, J. L., & Cepeda, G. 2018. Understanding relationship quality in hospitality services: A study based on text analytics and Partial Least Squares. Internet Research, 29 (3), 478-503. ˗ Satama, S. 2014. Consumer adoption of access-based consumption services. Case AirBnB. (Unpublished Master's Thesis). Aalto University, Finland. ˗ So, K.K.F., Oh, H. & Mina, S. 2018. Motivations and constraints of Airbnb consumers: Findings from a mixedmethods approach. Tourism Management, 67 (August), 224-236 ˗ Sparks, B. A., So, K. K. F. & Bradley, G. L. 2016. Responding to negative online reviews: The effects of hotel responses on customer inferences of trust and concern. Tourism Management, 53, 74–85. ˗ Stors, N. & Kagermeier, A. 2015. Motives for Using Airbnb in Metropolitan Tourism—Why do people sleep in the bed of a stranger? Regions Magazine, 299 (1), 17-19, ˗ Sundararajan, A. 2016. The sharing economy: The end of employment and the rise of crowd-based capitalism. Cambridge, MA: MIT Press. ˗ Taylor, P. 1995. Measuring changes in the relative competitiveness of package tour destinations. Tourism Economics, 1 (2), 169-82. ˗ Teubner, T., Hawlitschek, F. & Dann, D. 2017. Price determinants on Airbnb: How reputation pays off in the sharing economy. Journal of Self-Governance and Management Economics, 5 (4), 53–80. ˗ Thrane, C. 2007. Examining the determinants of room rates for hotels in capital cities: The Oslo experience. Journal of Revenue and Pricing Management, 5 (4), 315–323. ˗ Tussyadiah, I.P. & Pesonen, J. 2016. Drivers and barriers of peer-to-peer accommodation stay–an exploratory study with American and Finnish travellers. Current Issues in Tourism, 21 (6), 703-720. ˗ Tussyadiah, I.P. 2015. An exploratory study on drivers and deterrents of collaborative consumption in travel. In Tussyadiah, I. and Inversini, A. (eds.), Information and communication technologies in tourism 2015, Switzerland: Springer International Publishing. ˗ Tussyadiah, I.P. 2016. Factors of satisfaction and intention to use peer-to-peer accommodation, International Journal of Hospitality Management, 55, 70–80. ˗ Wang, D. & Nicolau, J.L. 2017. Price determinants of sharing economy based on accommodation rental: A study of listings from 33 cities on Airbnb.com. International Journal of Hospitality Management, 62 (April), 120–131. ˗ Xie, X. L. & Beni, G. 1991. A validity measure for fuzzy clustering. IEEE Transactions on Pattern Analysis & Machine Intelligence, (8), 841-847. ˗ Yang, S. & Ahn, S. 2016. Impact of motivation in the sharing economy and perceived security in attitude and loyalty toward Airbnb. Advanced Science and Technology Letters, 129, 180-184. 18 ˗ Zervas, G., Proserpio, D. & Byers, J.W. 2017. The rise of the sharing economy: estimating the impact of airbnb on the hotel industry, Journal of Marketing Research, 54, 687-705. ˗ Zhang, Z., Ye, Q. & Law, R. 2011. Determinants of hotel room price: An exploration of travellers’ hierarchy of accommodation needs. International Journal of Contemporary Hospitality Management, 23 (7), 972–981. ˗ Zhao, Y., Xu X. & Wang, M. 2019. Predicting overall customer satisfaction: Big data evidence from hotel online textual reviews. International Journal of Hospitality Management, 76 (Part A, January), 111-121.