Full text
A fuzzy approach for natural noise management in group recommender systems Jorge Castroa,b,∗ , Raciel Yerac , Luis Martínezd a Department of Computer Science and Artificial Intelligence, University of Granada, Calle Periodista Daniel Saucedo Aranda s/n, 18071 Granada, Spain b School of Software, University of Technology Sydney, PO Box 123, Broadway, Ultimo 2007 NSW, Australia c University of Ciego de Ávila, Carretera a Morón K m . 9 1/2, Ciego de Ávila, Cuba d Computer Science Department, University of Jaén, Campus L a s Lagunillas s/n, 23008 Jaén, Spain Abstract Information filtering is a k e y task in scenarios with information overload. Group Recommender Systems (GRSs) filter content regarding groups of users preferences and needs. Both the recommendation method and the available data influence recommendation quality. Most researchers improved group recommendations through the proposal of new algorithms. However, it has been pointed out that the ratings are not always right because users can introduce noise due to factors such as context of rating or user’s errors. This introduction of errors without malicious intentions is named natural noise, and it biases the recommendation. Researchers explored natural noise management in individual recommendation, but few explored it in GRSs. The latter ones apply crisp techniques, which results in a rigid management. In this work, w e propose Natural Noise Management for Groups based on Fuzzy Tools (NNMG-FT). NNMG-FT flexibilises the detection and correction of the natural noise to perform a better removal of natural noise influence in the ∗ Corresponding author Email addresses: [email protected] (Jorge Castro), [email protected] (Raciel Yera), [email protected] (Luis Martínez) Preprint submitted to Expert Systems with Applications December 16, 2017 Originally published in Expert Systems with Applications 94, 237-248, 2018
recommendation, hence, the recommendations of a latter GRS are then improved. Keywords: Natural noise,group recommender systems,collaborative filtering,fuzzy logic,computing with words 1. Introduction The Web allows people accessing to a huge amount of information. However, the users skills to cope with all the available information are limited, which leads to select suboptimal alternatives. This problem is known as information overload. Recommender Systems (RSs) are tools to help individuals to overcome such information overload problem personalizing access to information (Adomavicius & Tuzhilin, 2005; Ekstrand et al., 2011). However, some items tend to be consumed by groups of users, such as tourist attractions (Garcia et al., 2012) or television programmes (Said et al., 2011). With this purpose in mind, Group Recommender Systems (GRSs) (Masthoff, 2015) help groups of users to find suitable items according to their preferences and needs. Several techniques have been used to improve individual recommendation, such as neighborhood-based collaborative filtering (Sarwar et al., 2001), matrix factorisation (Koren et al., 2009), or approaches that consider temporal dynamics (Koren, 2010; Rafailidis et al., 2017). In the case of group recommendation, there are approaches to aggregate individual information (Masthoff, 2015), to consider consensus among members (Castro et al., 2015), or matrix factorisation models for groups (Ortega et al., 2016). A decade ago, it was pointed out that explicitly stated user preferences may not be error free (O’Mahony et al., 2006). More recently, other recent works (Bellogín et al., 2014; Centeno et al., 2015; Guo & Dunson, 2015; Zhang et al., 2
2017) have also pointed out that a person’s ratings are noisy, inconsistent, and biased. Li et al. (2013) determined that too many noisy ratings can distort users’ preference profiles, which result in unlike-minded neighbors that imply a quality loss in recommendations. Kluver et al. (2012) have also suggested that user ratings are imperfect and noisy, and such noise limits the predictive power of any RS. Therefore, in addition to improving recommendations through new recommendation approaches, researchers should also focus on improving the quality of the rating database (Amatriain et al., 2009c). In RSs, there are two kinds of noise in the database (O’Mahony et al., 2006): (i) malicious noise, that consists of erroneous data deliberately inserted in the system to influence recommendations, and (ii) natural noise, that appears when users unpurposely introduce erroneous data due to human errors or external factors during the rating process. This paper focuses on the latter. Natural noise biases recommendations, therefore, its management is a key factor to improve them. There are several Natural Noise Management (NNM) approaches for individual RSs databases. While some NNM approaches need additional information (Amatriain et al., 2009a; Pham & Jung, 2013), others detect and correct the natural noise using information already contained in the database (Yera et al., 2016; Yera Toledo et al., 2015). GRSs also rely on databases with explicit users’ preferences (Masthoff, 2015), therefore, they are affected by natural noise. Castro et al. (2017) propose a NNM approach for GRSs to manage ratings and noise using crisp values. This is the only work focused on NNM in GRSs. However, the crisp management is not either flexible or robust enough to deal with the uncertainty and vagueness of both the ratings and the NNM, which makes it necessary to develop new proposals with this regard. 3
In order to manage such uncertainty and vagueness in RSs contexts, the use of fuzzy tools has been considered for several years. A recent survey paper (Yera & Martínez, 2017) has shown that some traditional fuzzy tools have been successfully used for a more flexible and accurate information processing in RSs. However, it also shows that there are several research gaps related to the necessity of new fuzzy approaches focused on the use of emergent information sources and concentrated in new research trends in RSs. Specifically, the natural noise management (Martínez et al., 2016) is one of such research trends. Our purpose is to study the natural noise management in group recommendation with fuzzy tools. Therefore, in this work we propose Natural Noise Management for Groups based on Fuzzy Tools (NNMG-FT) to improve the rating database removing the natural noise. NNMG-FT applies three steps of management: fuzzy profiling, global noise management and local noise management. Both global natural noise management step and local noise management step are divided into two sub-steps: noise detection and noise correction. Both sub-steps apply fuzzy tools. In the noise detection, fuzzy tools allow to make a flexible classification of the ratings into noisy or not noisy. In the noise correction, this flexible classification is used to correct noisy ratings applying a soft modification of the value regarding its noise degree. The main advantages of NNMG-FT are: flexibility, robustness and consideration of group information in the NNM. A case study was performed to show the validity of NNMG-FT. In short, the main contributions of this paper consists of: •Design an improved profiling that manages uncertainty and vagueness of the ratings through the application of fuzzy tools in the profiling of ratings, users, and items. •Design an adequate representation of the noise management process that 4
improves the flexibility and robustness of the noise detection and noise correction. •Propose a NNM approach for GRSs that hybridizes several steps of noise detection and correction based on the information level from the viewpoints of both the whole ratings database and the groups ratings. •Validate the proposal through comparison with previous ones with similar purpose. The remainder of this paper is structured as follows. First, Section 2 presents the related works for the current research. Section 3 details NNMG-FT, our proposal for NNM in group recommendation. Section 4 shows the case study done to validate NNMG-FT performance. Finally, Section 5 concludes the work. 2. Related works In this section we revise different concepts about natural noise management in recommender systems, GRSs, and fuzzy sets, that are used in our NNM approach for GRSs. 2.1. Natural noise management The existence of underlying noise in users’ preferences in RSs and its negative effect have been referred for several years. In this way, an influential paper presented by Herlocker et al. (Herlocker et al., 2004) pointed out that, although an important amount of advanced algorithms were developed for improving RSs accuracy, the mean absolute error tends to yield around a constant magnitude. They speculated then that such algorithms could be reaching some magic barrier where natural variability in ratings may prevent researches from getting much more accurate results. The existence of such magic barrier has been confirmed 5
by further investigations in the last few years (Bellogín et al., 2014; Said et al., 2012), which have been focused on its characterisation and estimation. Additionally, the underlying noise in users’ preferences began to be referred as natural noise. Formally, natural noise term was first coined by O’Mahony et al. (2006) as those inconsistencies introduced in recommender systems databases due to the imperfect users behaviour when they rate the reviewed or purchased products, without a premeditated malicious intention. It is produced by the influence of external factors in the rating process, such as human errors or rating in different contexts. Natural noise influences the quality of user ratings, and researchers have determined that this influence results in poor recommendations (Amatriain et al., 2009b,c). Therefore, an adequate Natural Noise Management (NNM) is key to improve recommendations. Researchers have explored NNM for individual RSs, which is applied as a preprocessing stage done over the ratings database to reduce the impact of noisy information. Some techniques remove noisy information from the rating database, such as O’Mahony et al. (2006), which deletes both malicious and natural noisy ratings, or Li et al. (2013), which eliminates noisy but non malicious users. These works use the information already contained in the ratings database. However, they overlook important information from the dataset. There are works that rely on additional information to correct natural noisy ratings. Amatriain et al. (2009c) propose the mining and usage of a curated dataset with information provided by experts to reduce noise. Pham & Jung (2013) uses item attributes to build user models and correct ratings not matching the model, which is built using information of other users identified as experts. More recently, Bellogín et al. (2014) use item attributes 6
for measuring user coherence in recommender systems databases, showing that the recommendation performance is improved when less coherent users are discarded. Later, Yu et al. (2016) propose a correction approach for ratings associated to such less coherent users. Additionally, Saia et al. (2016) have presented an approach for removing incoherent items from a user profile, using semantic information. These approaches need additional information to correct the noisy ratings, which may not be feasible to obtain in certain domains. Recent proposals also focus on the detection and correction of natural noisy ratings using information contained in the original database. Some of these proposals use contradiction-based approaches (Yera Toledo et al., 2015) or fuzzy tools (Yera et al., 2016). On the other hand, in RSs context there are items that, because of their social features, tend to be consumed by groups, such as tour packages for groups of tourists (Ardissono et al., 2003), playlists for groups of listeners (Crossen et al., 2002), or healthy food for groups of family members, friends or colleagues (Trang Tran et al., 2017). In the social items scenario, Group Recommender Systems (GRSs, see next Section 2.2) have emerged as an effective solution for providing recommendations to groups of people (Castro et al., 2015; De Pessemier et al., 2014). Nevertheless, the NNM in GRSs has been an unexplored research area, regarding that most of the revised works focus on managing natural noise in individual RSs. The only reported research considering NNM in GRSs has been recently presented by Castro et al. (2017). Although the approaches proposed in such work introduced improvements in the recommendation accuracy, it also has some important shortcomings. Specifically, it does not consider the uncertainty and vagueness associated to the users preferences, which have been considered in NNM process for individual RSs (Yera et al., 2016). The latter NNM approach proved to 7
Table 1: Research works focused on natural noise management Target Individual Group NNM Crisp O’Mahony et al. (2006) Amatriain et al. (2009b,c) Li et al. (2013) Pham & Jung (2013) Bellogín et al. (2014) Saia et al. (2016) Yu et al. (2016) Yera Toledo et al. (2015) Castro et al. (2017) Fuzzy Yera et al. (2016) This contribution improve recommendation accuracy in comparison to previous crisp models for individual RSs (Li et al., 2013; Yera Toledo et al., 2015). Table 1 summarises the referred works and classifies them according to the recommendation context, individual or group, and to the techniques used for the NNM. Such table suggests that the NNM in groups using fuzzy techniques is still an area to explore. In this work, we aim to fill this gap with a new approach that, in contrast to the previous ones, performs an intensive use of fuzzy techniques both in the preferences of the active group and in the preferences of all available users, to perform a flexible and robust NNM. These features enhance the quality of the corrected users’ preferences and, therefore, improve the recommendations of the associated GRS. The magic barrier in accuracy (Bellogín et al., 2014) is not only caused by natural noise. Recommendation accuracy is also limited by temporal dynamics (Zhang et al., 2014), which study the evolution of users preferences across time due to changes in users taste. Different proposals have studied this important issue in RSs, such as time-based collective factorisation (Vaca et al., 2014), temporal matrix factorisation (Zhang et al., 2014) or combination of multimodal and temporal information (Rafailidis et al., 2017). Although temporal 8
dynamics is related to our research, it is based on another view of the experts’ preferences and it is not simultaneously considered with natural noise in this paper to avoid digressing and mix up its goal. 2.2. Group recommender systems GRSs provide groups of users with group personalised access to information (Masthoff, 2015). The group recommendation problem has been formalised, using the notation shown in Table 2, as finding the item, or set of items, that maximise the aggregated prediction for the target group: Recommendation(Ga,I) = argmax i∈IPrediction(Ga,i)(1) where Gais the target group, Iis the set of items in the database and Prediction(Ga,i)predicts the rating that group Gawould give to item iand is given in the [rmin,rmax]domain. Table 2: Notation used for group recommender systems. Symbol Description U={u1,...,um}set of all users I={i1,...,in}set of all items R⊆U×Iset of known ratings. rui ∈Rrating that user ugave about item i [rmin,rmax]rating domain of the dataset given between rmin and rmax Ru⊆Rset of ratings given by user u. Ri⊆Rset of ratings for item i. Ga={m1,...,mg} ⊆ Utarget group with gmembers RGa i⊆Riratings that members of group Gaprovided for item i. Among the various ways to compute the prediction for the group, the most successful ones are based on aggregating individual information to recommend: •Rating aggregation (Kagita et al., 2015): The rating profiles of all members are aggregated into a single rating profile that represents the group preferences. This group profile is then used in an individual recommender system to recommend. 9
detect clear tendencies. With this regard, the fuzzy profiles are modified using a soft modification that can be formulated in various ways depending of the aim. The proposal uses transformation function f1, presented in Fig. 3, which depends on parameter kthat indicates the extent to which unclear tendencies are attenuated. The best value for parameter kis determined through an empirical analysis. 0 1 01 k f1 Modi ed membership degree Original membership degree Figure 3: The fuzzy transformation function. Therefore, Eq. 10, 11, 12, and 13 present the transformed user, item and rating profiles, which are used in the following phases. p∗ u=f1(pulow ),f1(pumedium ),f1(puhigh )(10) p∗ i=f1(pilow ),f1(pimedium ),f1(pihigh )(11) p∗Ga i=f1(pGa ilow ),f1(pGa imedium ),f1(pGa ihigh )(12) p∗ rui =f1(µlow(rui)),f1(µmedium(rui)),f1(µhigh(rui))(13) 3.2. Global noise management Once the fuzzy profiles are obtained, the global rating correction phase is performed, which is presented in this section. This phase aims to perform an initial reduction of the natural noise from the viewpoint of the entire ratings 16
database. To do so, this phase is divided into two steps: (a) global noise detection, which detects noisy ratings, and (b) global noise correction, which modifies the noisy ratings value to reduce the natural noise in the ratings database. The inputs of this phase are the rating dataset and the fuzzy profiles obtained in the previous phase. As output, this phase produces a de-noised ratings database, which will be used in the third phase. 3.2.1. Global noise detection This step develops an exhaustive analysis of each rating to find noisy ones, therefore its inputs are the ratings database and the fuzzy profiles, and its output is a list of detected noisy ratings. Specifically, the aim is to find ratings whose corresponding user and item have consistent rating tendencies and the rating itself is not coherent with them because this situation might indicate that the rating value is noisy. Eq. 14 formalises this strategy to decide whether the current rating rui is noisy or not using these two conditions. The first condition evaluates, for a rating rui, whether its corresponding user fuzzy profile p∗ uand item fuzzy profile p∗ ihave similar preference tendencies checking whether they are close enough. The second condition evaluates whether the rating fuzzy profile p∗ rui is far enough from both the user and the item fuzzy profiles. It means that the rating value does not match its corresponding user and item tendencies. If the rating rui satisfies both conditions, then it is considered as noisy. (d(p∗ u,p∗ i)< δ1 | {z } first condition and min(d(p∗ u,p∗ rui ), d(p∗ i,p∗ rui )) >=δ2 | {z } second condition )→rui is noisy (14) Eq. 14 depends on a dissimilarity function d. In this proposal, the Manhattan distance is used as dissimilarity measure (see Eq. 15) because it 17
reflects the differences between dimensions without giving importance to how these differences are distributed among dimensions, as euclidean distance would do. The best values for δ1and δ2are determined through an empirical analysis. d(p∗ u,p∗ i) = X s |p∗ us−p∗ is| =|p∗ ulow −p∗ ilow |+|p∗ umedium −p∗ imedium |+|p∗ uhigh −p∗ ihigh |(15) 3.2.2. Global noise correction Once the noisy ratings have been identified, this step aims at correcting them. This step receives as input the list of detected noisy ratings, and performs a flexible correction of the noisy ratings whose output is a de-noised ratings database. To do so, noisy ratings are characterised by their noise degree, which controls the extent of the correction. This noisy degree is computed using the fuzzy profiles used in the previous phases. The formal definition of the noise degree is given by Eq. 16, which is computed considering the minimum dissimilarity between the rating, and user or item fuzzy profiles, i.e., the value of the second condition in Eq. 14. The Manhattan distance between fuzzy profiles is restricted to the interval [1, 2], as it was proved in (Yera et al., 2016). Thus, the noise degree is normalised subtracting 1 to the minimum distance. NoiseDegreerui =min(d(p∗ u,p∗ rui ), d(p∗ i,p∗ rui ))−1(16) The de-noised rating value r∗ ui is computed through a convex combination of the rating value rui and the prediction nui, which is controlled by the noise 18
degree. Eq. 17 formalises this procedure, which is applied to all noisy ratings. r∗ ui =rui ∗(1−NoiseDegreerui )+ nui ∗NoiseDegreerui (17) where nui is a predicted rating value that is obtained using a collaborative filtering rating prediction approach that considers all the available rating data. We propose to use the user-based collaborative filtering approach (Ning et al., 2015), as previous researches used (Li et al., 2013), although other approaches could be used. nui =Prediction(u,i,R)∈[rmin,rmax](18) As a result of this phase, the corrected ratings dataset R∗is obtained, which is used as input in the following phase. 3.3. Local noise management Once the global noise management phase is completed, the last phase applies a similar NNM focused in the local level, i.e., focused on the group ratings. To do so, this phase takes as inputs the database R∗already de-noised by phase 2 and the corresponding fuzzy profiles. It refines R∗with a NNM focused on the target group ratings, which generates a new ratings database R∗∗ as output. This phase is also composed of two steps: (i) local noise detection and (ii) local noise correction. 3.3.1. Local noise detection In the case of the local noise detection, all the ratings of group Gamembers received as input, are checked to verify whether they are natural noisy. Eq. 19 presents the criteria used to detect whether rating rui is noisy using in this case the group-based item profile (see Eq. 9). This step produces a list of noisy 19
ratings as output. (d(p∗ u,pGa∗ i)< δ1and min(d(p∗ u,p∗ rui ), d(pGa∗ i,p∗ rui )) >=δ2)→rui is noisy (19) 3.3.2. Local noise correction The local noise correction step corrects all the ratings identified as noisy in the local noise detection step (input), and produces a de-noised ratings database (output). Here, similarly to the previous phase, the group-based fuzzy item profile is considered. Eq. 20 formalises the calculation of the noise degree regarding the group Ga. NoiseDegreeGa r∗ ui =min(d(p∗ u,p∗ rui ), d(pGa∗ i,p∗ rui ))−1(20) The noise correction is performed then over the noisy ratings using Eq. 21. As a result, the noise in the rating database is reduced with a correction adjusted to the target group. r∗∗ ui =r∗ ui ∗(1−NoiseDegreeGa r∗ ui )+n∗ ui ∗NoiseDegreeGa r∗ ui (21) where n∗ ui is a predicted rating value that is obtained from the corrected rating database R∗. Similarly to the computation of nui, we propose to use the userbased collaborative filtering approach (Ning et al., 2015). n∗ ui =Prediction(u,i,R∗)∈[rmin,rmax](22) This last step concludes the third phase and, therefore, NNMG-FT. The output is a de-noised rating database R∗∗ that can later be used by a GRS to 20
recommend. 4. Case study We developed a case study to evaluate NNMG-FT. The remaining of this section details the experimental protocol and shows its results. 4.1. Experimental protocol To evaluate NNMG-FT, we used an experimental procedure based on a popular protocol for group recommendation (De Pessemier et al., 2014). Such experimental procedure is composed of the following steps: •Partition the rating dataset in training and test sets randomly. •Generate the groups randomly. •Apply the proposed NNM approach to the training set. •Recommend to each group regarding the data in the training set and the group recommendation algorithm. •Evaluate the recommendations using the test set. Within this experimental protocol, the Mean Absolute Error has been considered as the evaluation metric. Specifically, such protocol was repeated 20 times and the values obtained in each execution were averaged. In each execution, 50 different random groups were generated to evaluate their recommendations. Groups sizes 5, 10, and 15 were considered in the case study. In this evaluation, three NNM approaches were compared: (i) Base, (ii) NNMG-Crisp, crisp NNM for groups (Castro et al., 2017), and (iii) the current proposal NNMG-FT, NNM for groups based on fuzzy tools. To quantify the 21
effect of each NNM approach, we measured the MAE of various GRSs with de-noised ratings databases using these NNM approaches. To perform a comprehensive analysis, we evaluated them with GRSs based on rating and recommendation aggregation approaches (De Pessemier et al., 2014). Various aggregation strategies can be applied within each of these aggregation approaches. Specifically, average and least misery strategies were used in the case study. These GRSs rely on a individual RS to recommend. All evaluated GRSs use the item-based collaborative filtering approach (Sarwar et al., 2001). This approach has been a very popular collaborative filtering method whose importance is high in the RSs field due to its simplicity, effectiveness and scalability (Adomavicius & Tuzhilin, 2005; Ekstrand et al., 2011). The case study comprises the evaluation in two well-known recommendation datasets: •The MovieLens 100K dataset1, which was collected by GroupLens Research Project at the University of Minnesota. It is composed of 100,000 ratings given by 943 users over 1,682 movies in the five stars domain. •The Netflix Tiny dataset, composed of 4,427 users, 1,000 movies, and 56,136 ratings, which were also given in the five stars domain. This is a smaller version of Netflix dataset, and it is available in the Personalised Recommendation Algorithms Toolkit 2. 1http://grouplens.org/datasets/movielens/100k/ 2http://prea.gatech.edu 22
4.2. Results The results of the experimental procedure are presented in this section. First, we perform an optimisation of the parameters of NNMG-FT. After that, we analyse the results of the considered NNM approaches for recommendation aggregation GRSs and for rating aggregation GRSs. Finally, MAE improvement per group of NNMG-FT is analysed. 4.2.1. Parameter optimisation of the NNMG-FT NNMG-FT has three parameters whose values need to be determined to adjust it. These parameters are k,δ1and δ2, and their values are determined through experiments to determine the best value for each of them. Parameter kranges from 0 to 1 and it is used to determine how fuzzy profiles of users and items are modified to highlight rating tendencies. Table 5 shows the MAE of NNMG-FT in various group recommendation scenarios. The whole range of values for kwas evaluated, here only values 0.35, 0.50 and 0.75 are shown for the sake of clearness. The results determine that NNMG-FT obtains the best MAE for k= 0.35. Parameter δ1ranges from 0 to 2 (Yera et al., 2016), and the larger its value the more close a user and item profile have to be in order to consider them as matching tendencies. Table 6 shows the MAE of NNMG-FT in various configurations for dataset, aggregation approach, aggregation strategy and group size. The whole range of values for δ1was evaluated, here only values 0.9, 1.0 and 1.1 are shown for the sake of clearness. The results show that NNMG-FT obtains the best outcomes for δ1= 1.0in most of the evaluated scenarios. Parameter δ2range from 0 to 2 (Yera et al., 2016), and the larger its value the more a rating has to deviate from its corresponding user-item tendency to be considered as noisy. Table 7 shows the MAE of NNMG-FT in various 23
Table 5: NNMG-FT parameter optimisation. Optimisation of kusing MAE. Aggregation Dataset Aggregation Group k approach strategy size 0.35 0.50 0.75 50.8364 0.8503 0.8652 Mean 10 0.8590 0.8684 0.8883 MovieLens 15 0.8686 0.8898 0.9093 100k 5 0.8545 0.8615 0.8864 Min 10 0.9177 0.9301 0.9544 Rating 15 0.9805 0.9962 1.0247 aggregation 5 0.8272 0.8416 0.8612 Mean 10 0.8517 0.8685 0.8857 Netflix 15 0.8566 0.8698 0.8907 Tiny 5 0.8555 0.8730 0.8891 Min 10 0.9167 0.9327 0.9520 15 0.9595 0.9703 0.9891 50.8335 0.8497 0.8791 Mean 10 0.8590 0.8726 0.8987 MovieLens 15 0.8802 0.8980 0.9110 100k 5 0.9855 1.0115 1.0900 Min 10 1.1050 1.1284 1.1402 Recomm. 15 1.1690 1.1840 1.2086 aggregation 5 0.8502 0.8725 0.8915 Mean 10 0.8335 0.8522 0.8621 Netflix 15 0.8528 0.8731 0.8927 Tiny 5 0.9888 0.9991 1.0152 Min 10 0.9855 0.9974 1.0130 15 1.1857 1.1965 1.2192 configurations. The whole range of values for δ2was evaluated, here only values 0.9, 1.0 and 1.1 are shown for the sake of clearness. The results show that NNMG-FT obtains the best outcomes for δ2= 1.0in most of the evaluated scenarios. The parameter optimisation results determine that the best configuration for NNMG-FT parameters are k= 0.35,δ1= 1 and δ2= 1. The remaining experiments are performed with those values. 4.2.2. Noise management in recommendation aggregation GRSs Table 8 shows the MAE of the recommendation aggregation GRSs with the compared NNM approaches. The lower the MAE of the GRS, the better the NNM approach. The best NNM of each configuration is highlighted in bold. 24
Table 6: NNMG-FT parameter optimisation. Optimisation of δ1using MAE. Aggregation Dataset Aggregation Group δ1 approach strategy size 0.9 1.0 1.1 5 0.8367 0.8364 0.8365 Mean 10 0.8595 0.8590 0.8591 MovieLens 15 0.8690 0.8686 0.8687 100k 5 0.8545 0.8545 0.8547 Min 10 0.9179 0.9177 0.9183 Rating 15 0.9808 0.9805 0.9810 aggregation 5 0.8278 0.8272 0.8274 Mean 10 0.8520 0.8517 0.8517 Netflix 15 0.8569 0.8566 0.8567 Tiny 5 0.8558 0.8555 0.8552 Min 10 0.9172 0.9167 0.9166 15 0.9599 0.9595 0.9594 5 0.8340 0.8335 0.8340 Mean 10 0.8595 0.8590 0.8591 MovieLens 15 0.8804 0.8802 0.8802 100k 5 0.9860 0.9855 0.9857 Min 10 1.1055 1.1050 1.1061 Recomm. 15 1.1694 1.1690 1.1693 aggregation 5 0.8504 0.8502 0.8502 Mean 10 0.8340 0.8335 0.8340 Netflix 15 0.8532 0.8528 0.8528 Tiny 5 0.9885 0.9888 0.9888 Min 10 0.9860 0.9855 0.9857 15 1.1866 1.1857 1.1860 NNMG-FT achieved the best results as compared with other approaches in all evaluated scenarios. Beyond this general improvement, there were differences regarding the relative improvement across datasets, aggregation strategies and group sizes. Table 9 shows, for the various configurations of datasets, aggregation strategies and group sizes, the relative improvement of the NNM approaches. Additionally, for the comparison between NNM-Crisp and NNM-FT, it is shown both the relative improvement and the p-value of the Wilcoxon signed-rank test (significant values with α= 0.05 are highlighted). The relative improvement has been calculated dividing the MAE of the first technique by the MAE of the reference technique. Wilcoxon test has been performed 25
Table 11: Relative improvement of the pairwise comparison of NNMG approaches on rating aggregation. Note that, for the comparison of NNM-Crisp and NNM-FT, the p-value of Wilcoxon signed-rank test is shown (statistically significant values with α= 0.05 are highlighted). Dataset Aggregation Group NNMG-Crisp NNMG-FT NNMG-Crisp strategy size vs Base vs Base vs NNMG-FT Rel. imp. p-value Avg 5 2.20% 3.43% 1.25% <0.001 10 2.09% 3.41% 1.34% <0.001 MovieLens 15 2.15% 3.34% 1.22% <0.001 100k Min 5 2.93% 3.74% 0.84% <0.001 10 3.83% 4.02% 0.20% 0.006 15 4.70% 4.33% -0.39% <0.001 Avg 5 0.68% 1.24% 0.56% <0.001 10 0.82% 1.35% 0.54% <0.001 Netflix 15 0.90% 1.30% 0.41% <0.001 Tiny Min 5 1.20% 1.51% 0.31% 0.016 10 2.01% 1.91% -0.10% 0.162 15 2.50% 2.14% -0.37% 0.001 5. Conclusions This paper proposes a natural noise management approach for group recommender systems using fuzzy tools (NNMG-FT). Specifically, NNMG-FT uses fuzzy profiles to characterise the rating tendency of users and items. With this characterisation, ratings that do not follow their corresponding user and item tendency are identified as noisy and, therefore, corrected. NNMG-FT performs two phases of noise correction: the first one follows a global approach, and the second is personalised to the target group. A case study has been performed to compare NNMG-FT with previous natural noise management approaches. The results show that the management of natural noise with our proposal leads to improved results in the majority of evaluation scenarios, which comprise various aggregation approaches, aggregation strategies and group sizes. Moreover, a deeper study of the proposal showed that the improvement of recommendations is general and few groups had a decay in recommendation quality. 32
Figure 8: MAE improvement per group shown as percentile for recommendation aggregation GRS. (a) Mean aggr. on MovieLens 100k. (b) Min. aggr. on MovieLens 100k. (c) Mean aggr. on Netflix Tiny. (d) Min. aggr. on Netflix Tiny. The study shows that NNMG-FT is beneficial for group recommendation. In order to further improve the NNM in future works, it is worth to study temporal dynamics, which enhance user preference modelling. Consideration of temporal dynamics would help at both detecting more noisy ratings and avoiding false positives, and therefore improve the detection of noise. Future works will also focus on exploring NNM in context-aware scenarios. Context in recommender systems is characterised by its heterogeneity, covering very diverse information sources, such as temporal information, companion, or weather. Moreover, context-awareness leads to a higher sparsity of ratings. Therefore, specific researches are needed to study the particularities of contextaware scenario, in order to characterise natural noise in group recommender systems databases. 33
Figure 9: MAE improvement per group shown as percentile for rating aggregation GRS. (a) Mean aggr. on MovieLens 100k. (b) Min. aggr. on MovieLens 100k. (c) Mean aggr. on Netflix Tiny. (d) Min. aggr. on Netflix Tiny. Acknowledgements This paper was partially supported by the Spanish FPU fellowship (FPU13/01151), the Spanish National research project TIN2015-66524-P. References Adomavicius, G., & Tuzhilin, A. T. (2005). Toward the next generation of recommender systems: A survey of the state-of-the-art and possible extensions. IEEE Transactions on Knowledge and Data Engineering,17, 734–749. Amatriain, X., Lathia, N., Pujol, J. M., Kwak, H., & Oliver, N. (2009a). The wisdom of the few: a collaborative filtering approach based on expert opinions 34
from the web. In 32nd International ACM SIGIR Conference (pp. 532–539). New York, NY, USA: ACM. Amatriain, X., Pujol, J. M., & Oliver, N. (2009b). I like it... i like it not: Evaluating user ratings noise in recommender systems. In User modeling, adaptation, and personalization (pp. 247–258). Springer. Amatriain, X., Pujol, J. M., Tintarev, N., & Oliver, N. (2009c). Rate it again: increasing recommendation accuracy by user re-rating. In third ACM conference on Recommender systems (pp. 173–180). ACM. Ardissono, L., Goy, A., Petrone, G., Segnan, M., & Torasso, P. (2003). Intrigue: Personalized recommendation of tourist attractions for desktop and hand held devices. Applied Artificial Intelligence,17, 687–714. Bellogín, A., Said, A., & de Vries, A. P. (2014). The magic barrier of recommender systems–no magic, just ratings. In International Conference on User Modeling, Adaptation, and Personalization (pp. 25–36). Springer. Castro, J., Quesada, F. J., Palomares, I., & Martínez, L. (2015). A consensusdriven group recommender system. International Journal of Intelligent Systems,30, 887–906. Castro, J., Yera, R., & Martínez, L. (2017). An empirical study of natural noise management in group recommendation systems. Decision Support Systems, 94, 1 – 11. Centeno, R., Hermoso, R., & Fasli, M. (2015). On the inaccuracy of numerical ratings: dealing with biased opinions in social networks. Information Systems Frontiers,17, 809–825. Crossen, A., Budzik, J., & Hammond, K. J. (2002). Flytrap: Intelligent group music recommendation. In Proceedings of the 7th International Conference 35
on Intelligent User Interfaces IUI ’02 (pp. 184–185). New York, NY, USA: ACM. De Pessemier, T., Dooms, S., & Martens, L. (2014). Comparison of group recommendation algorithms. Multimedia Tools and Applications,72, 2497– 2541. Ekstrand, M. D., Riedl, J. T., & Konstan, J. A. (2011). Collaborative filtering recommender systems. Foundations and Trends in Human-Computer Interaction,4, 81–173. Garcia, I., Pajares, S., Sebastia, L., & Onaindia, E. (2012). Preference elicitation techniques for group recommender systems. Information Sciences,189, 155 – 175. Guo, F., & Dunson, D. B. (2015). Uncovering systematic bias in ratings across categories: a bayesian approach. In Proceedings of the 9th ACM Conference on Recommender Systems (pp. 317–320). ACM. Herlocker, J., Konstan, J., Terveen, K., & Riedl, J. (2004). Evaluating collaborative filtering recommender systems. ACM Transactions on Information Systems,22, 5–53. Herrera, F., & Martinez, L. (2000). A 2-tuple fuzzy linguistic representation model for computing with words. IEEE Transactions on Fuzzy Systems,8, 746–752. Kagita, V. R., Pujari, A. K., & Padmanabhan, V. (2015). Virtual user approach for group recommender systems using precedence relations. Information Sciences,294 , 15 – 30. Innovative Applications of Artificial Neural Networks in Engineering. 36
Kluver, D., Nguyen, T. T., Ekstrand, M., Sen, S., & Riedl, J. (2012). How many bits per rating? In Proceedings of the sixth ACM conference on Recommender systems (pp. 99–106). ACM. Koren, Y. (2010). Collaborative Filtering with Temporal Dynamics. Interacting with Computers OF THE ACM,53, 89–97. Koren, Y., Bell, R., & Volinsky, C. (2009). Matrix factorization techniques for recommender systems. Computer,42 , 30–37. Li, B., Chen, L., Zhu, X., & Zhang, C. (2013). Noisy but non-malicious user detection in social recommender systems. World Wide Web,16, 677–699. Martínez, L., Castro, J., & Yera, R. (2016). Managing natural noise in recommender systems. In C. Martín-Vide, T. Mizuki, & M. A. Vega-Rodríguez (Eds.), Theory and Practice of Natural Computing: 5th International Conference, TPNC 2016, Sendai, Japan, December 12-13, 2016, Proceedings (pp. 3–17). Springer International Publishing. Masthoff, J. (2015). Group recommender systems: Aggregation, satisfaction and group attributes. In F. Ricci, L. Rokach, & B. Shapira (Eds.), Recommender Systems Handbook (pp. 743–776). Springer US. McCarthy, J. F., & Anagnost, T. D. (1998). Musicfx: an arbiter of group preferences for computer supported collaborative workouts. In The 1998 ACM Conference on Computer Supported Cooperative Work (pp. 363–372). ACM. Ning, X., Desrosiers, C., & Karypis, G. (2015). A comprehensive survey of neighborhood-based recommendation methods. In F. Ricci, L. Rokach, & B. Shapira (Eds.), Recommender Systems Handbook (pp. 37–76). Springer US. 37
O’Connor, M., Cosley, D., Konstan, J. A., & Riedl, J. (2001). Polylens: a recommender system for groups of users. In Proceeding of the 7th Conference on European Conference on Computer Supported Cooperative Work (ECSCW’01) (pp. 199–218). Bonn, Germany – September 16-20. O’Mahony, M. P., Hurley, N. J., & Silvestre, G. (2006). Detecting noise in recommender system databases. In 11th international conference on Intelligent user interfaces (pp. 109–115). ACM. Ortega, F., Hernando, A., Bobadilla, J., & Kang, J. H. (2016). Recommending items to group of users using matrix factorization based collaborative filtering. Information Sciences,345 , 313 – 324. Pham, H. X., & Jung, J. J. (2013). Preference-based user rating correction process for interactive recommendation systems. Multimedia tools and applications,65 , 119–132. Rafailidis, D., Kefalas, P., & Manolopoulos, Y. (2017). Preference dynamics with multimodal user-item interactions in social media recommendation. Expert Systems with Applications,74, 11 – 18. Rodríguez, R. M., Labella, A., & Martínez, L. (2016). An overview on fuzzy modelling of complex linguistic preferences in decision making. International Journal of Computational Intelligence Systems,9, 81–94. Rodríguez, R. M., & Martínez, L. (2013). An analysis of symbolic linguistic computing models in decision making. International Journal of General Systems,42, 121–136. Saia, R., Boratto, L., & Carta, S. (2016). A semantic approach to remove incoherent items from a user profile and improve the accuracy of a 38
recommender system. Journal of Intelligent Information Systems,47, 111– 134. Said, A., Berkovsky, S., & De Luca, E. W. (2011). Group recommendation in context. In 2nd Challenge on Context-Aware Movie Recommendation CAMRa ’11 (pp. 2–4). New York, NY, USA: ACM. Said, A., Jain, B. J., Narr, S., & Plumbaum, T. (2012). Users and noise: The magic barrier of recommender systems. In J. Masthoff, B. Mobasher, M. C. Desmarais, & R. Nkambou (Eds.), User Modeling, Adaptation, and Personalization: 20th International Conference, UMAP 2012, Montreal, Canada, July 16-20, 2012. Proceedings (pp. 237–248). Springer Berlin Heidelberg. Sarwar, B., Karypis, G., Konstan, J., & Riedl, J. (2001). Item-based collaborative filtering recommendation algorithms. In 10th international conference on World Wide Web (pp. 285–295). ACM. Schweizer, B., & Sklar, A. (2011). Probabilistic metric spaces. Courier Corporation. Trang Tran, T. N., Atas, M., Felfernig, A., & Stettinger, M. (2017). An overview of recommender systems in the healthy food domain. Journal of Intelligent Information Systems, . Vaca, C. K., Mantrach, A., Jaimes, A., & Saerens, M. (2014). A timebased collective factorization for topic discovery and monitoring in news. In Proceedings of the 23rd International Conference on World Wide Web WWW ’14 (pp. 527–538). New York, NY, USA: ACM. Yager, R. R., & Zadeh, L. A. (2012). An introduction to fuzzy logic applications in intelligent systems volume 165. Springer Science & Business Media. 39
Yera, R., Castro, J., & Martínez, L. (2016). A fuzzy model for managing natural noise in recommender systems. Applied Soft Computing,40, 187 – 198. Yera, R., & Martínez, L. (2017). Fuzzy tools in recommender systems: A survey. International Journal of Computational Intelligence Systems,10, 776–803. Yera Toledo, R., Caballero Mota, Y., & Martínez, L. (2015). Correcting noisy ratings in collaborative recommender systems. Knowledge-Based Systems,76, 96 – 108. Yu, P., Lin, L., & Yao, Y. (2016). A novel framework to process the quantity and quality of user behavior data in recommender systems. In B. Cui, N. Zhang, J. Xu, X. Lian, & D. Liu (Eds.), Web-Age Information Management: 17th International Conference, WAIM 2016, Nanchang, China, June 3-5, 2016, Proceedings, Part I (pp. 231–243). Springer International Publishing. Yu, Z., Zhou, X., Hao, Y., & Gu, J. (2006). Tv program recommendation for multiple viewers based on user profile merging. User Modeling and UserAdapted Interaction,16, 63–82. Zadeh, L. A. (1965). Fuzzy sets. Information and control,8, 338–353. Zadeh, L. A. (2012). Computing with Words: Principal Concepts and Ideas. Studies in Fuzziness and Soft Computing. Springer. Zhang, C., Wang, K., Yu, H., Sun, J., & Lim, E.-P. (2014). Latent factor transition for dynamic collaborative filtering. In Proceedings of the 2014 SIAM International Conference on Data Mining (pp. 452–460). SIAM. Zhang, X., Zhao, J., & Lui, J. (2017). Modeling the assimilation-contrast effects in online product rating systems: Debiasing and recommendations. In Proceedings of the Eleventh ACM Conference on Recommender Systems (pp. 98–106). ACM. 40
Zimmermann, H.-J. (2001). Fuzzy set theory—and its applications. Springer Science & Business Media. 41