scieee AI-readable full text Open interactive document viewer

A Machine Learning Model for a System that Recommends Movies

Annual Methodological Archive Research Review (AMARR)

Full text

http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 466 A Machine Learning Model for a System that Recommends Movies Zeeshan Ali Department of Computer Science, Minhaj University, Lahore, Pakistan Email: [email protected] Hafiz Muhammad Zubair Afzal Department of Computer Science, Minhaj University, Lahore, Pakistan Email: [email protected] Muhammad Yousuf* Department of Computer Science, Minhaj University, Lahore, Pakistan Email: [email protected] Sohail Ahmad Department of Computer Science, Minhaj University, Lahore, Pakistan Email: [email protected] Sonia Mukhtar Department of Computer Science, Minhaj University, Lahore, Pakistan Email: [email protected] Recommendation systems primarily aim to offer customers useful product suggestions by relying solely on past interactions. Recommender systems stand out as particularly useful in businesses due to their application of machine learning technologies. This form of recommendation filtering is used to attempt to predict a user’s selection. With the help of data, it forecasts, aims, and even identifies what the consumers’ needs are from an ever-growing assortment of options. Multiple markers such as a user’s search history, their age and background, what they have bought previously, and a lot more, can help locate the users. It helps users locate products and services which are unavailable or difficult for them to find. People now find it difficult to locate and sort through their preferred content due to the deluge of information. This issue has been addressed by recommendation systems (RSs); However, there are significant problems with data scalability, data scarcity, and the cold-start problem with traditional Appen recommendation systems, like contentbased and collaborative filtering, all of which require complex solutions.Data sparsity and a failure to consider the variety of recommended outcomes are two issues with traditional recommendation systems. While the second experiment extended predictions to 4800 movies and produced a SVM 86% accuracy as compared to others. Introduction These systems are now widely employed in many different industries, including movies, music, books, videos, apparel, restaurants, food, locations, and many more. http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 467 They have also grown increasingly stylish. Finding stuff that would be exciting to a person is the main goal of a recommendation system. Additionally, it has several features to generate personalized lists of interesting and helpful stuff for each user or person. Recommendation systems have many different applications. These are now used in the majority of our internet platforms and have grown in popularity over the past few years. These platforms contain a wide range of content, including movies, music, books, and videos; friends and stories on social media; products on e-commerce websites; people on dating and professional websites; and Google search results. Over time, recommendation systems have been able to grow more complex thanks to the advent of big data and machine learning. In the future, recommendation systems may be able to handle large amounts of data and use state-of-the-art methods like deep learning and reinforcement learning to improve the quality of their recommendations and make them more relevant to the data's user [1]. Due to the use of implicit data in these systems, there has been a recent surge in research on cognitive-based recommendation systems that analyze users' personalities or behaviors to ascertain their preferences. This method has the advantage of enabling recommendation systems to promptly adjust to shifts in user preferences. [2]. Stepwise is the main technique employed in the movie or any recommender system (RS), where users rate a number of items before the system forecasts their ratings for an item that hasn't been reviewed yet. [3] By creating personalized lists of goods, resources, and data based on consumers' preferences and prior experiences, they aid in decision-making. Furthermore, RSs employ a variety of technologies to weed out results that are satisfactory, cutting down on search time and delivering the information consumers require fast. Social networking, entertainment, e-commerce, scientific research, news, healthcare, tourism, and education are just a few of the industries that have seen a sharp increase in the use of RS in recent years. [4] Personalized recommendation systems mine user behavior data and recommend products that users are likely to find highly interesting using artificial intelligence, data mining, and other related Internet technologies [5]. For example, RSs algorithms are used by YouTube and Netflix to provide tailored video and content recommendations. These suggestions are mostly derived from content analytics, user behaviour trends, and social connections. In order to suggest material that fits with users' interests and previous interactions, the systems leverage commonalities in friend relationships and listening histories. This approach keeps users hooked on the platform by ensuring they consistently find fresh, interesting information that suits their interests. In a similar vein, RS algorithms are used by social networking and content-sharing businesses such as Facebook, Instagram, and LinkedIn to increase user engagement. In order to increase efficiency and the user experience overall, our study model shows how to make use of the enormous number of movies that are available online to suggest movies that people are likely to appreciate. Users can find superior material to watch more easily as a result. In line with the continuous developments in streaming services and digital entrepreneurship, the study offers a personalized movie recommendation model powered by AI that incorporates sentiment and content analysis. The model tackles the issues of content overload by incorporating AI technology, improving the digital entertainment environment for both consumers and service providers. These days, a wide range of industries, including entertainment and education, use recommendation http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 468 algorithms. Customers used to have to choose which movies to watch, what music to listen to, which books to purchase, and so on. Commercial movie libraries have more than 15 million films, far more than any one person could ever watch. The sheer number of films available for viewing can be overwhelming. Consequently, both movie service providers' and consumers' enthusiasm depend on an effective recommendation system. [6] The recommendation machine has gained popularity as e-commerce and the Internet have developed. The collaborative filtering technique is examined similarly in this paper's electronic commerce recommendation system, which specializes in the use of personalized movie recommendation systems [7]. To use both explicit and implicit user data for item suggestion, a range of data mining techniques must be used to understand a user's preferences. As a result, there is continual study on methods to improve the insights provided for recommendations based on previously chosen items, user input on recommendation outcomes, user correlation analysis, etc. [8–13]. For programming purposes, its functions make use of a resilient distributed dataset (RDD). Nonetheless, the ability to run programs in parallel could be a strength. Scala, Python, and R are just a few of the libraries and tools that Spark offers to help manage and scale data [14]. People with similar tastes are given recommendations for items. Lastly, the hybrid approach enhances the quality of its suggestions by utilizing content-based filtering, collaborative methods, and other tactics [15]. Literature Review: Movie recommendation systems have become a major area of machine learning application, aiming to help users navigate vast movie catalogs by predicting what they will likely enjoy. Previously reliant on rule-driven and popularity-based approaches, the emergence of large datasets such as MovieLens and the Netflix Prize accelerated the adoption of data-driven techniques. Conventional approaches include contentbased filtering, which represents movies using features like genre, cast, and plot keywords and recommends similar items to those a user has liked, and collaborative filtering, which uses either memory-based similarity models or model-based techniques like matrix factorization and SVD++ to learn patterns from user-item interactions. Content-based approaches are good at handling new items, but they can have a limited scope, and collaborative filtering has issues with cold-start and sparsity. To increase accuracy and robustness, hybrid systems combine the two by employing switching techniques, score blending, or feature augmentation. Through neural collaborative filtering, autoencoders, and sequence models like RNNs, GRUs, and Transformers (e.g., SASRec, BERT4Rec) that capture temporal dynamics of user preferences, deep learning has revolutionized the field in recent years. By combining textual elements from plots and reviews, visual elements from posters, and even audio from trailers which is particularly helpful in cold-start situations multimodal learning has further enhanced recommendations. To increase the accuracy of recommendations, graph-based techniques such as knowledge graph reasoning and graph neural networks take advantage of the relationships between users, actors, directors, and films. Privacy-maintaining Collaborative filtering has been gaining more attention as a result of the growing requirement to keep private data while also providing advice. Several methods were put forth to estimate pointers without presenting a serious privacy risk in order to improve the comfort of statistics owners' http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 469 experiences while still making predictions.By using outstanding privacy-preserving techniques, such methods eliminate or lessen the privacy, financial, and legal concerns of statistics owners [16]. The recommendation machine has gained popularity as e-commerce and the Internet have developed. This paper examines the collaborative filtering algorithm in the context of a personalized film recommendation system, with a similar focus on electronic commerce [17]. In this paper, we learn about the KNN algorithm and collaborative filtering algorithm to design and implement a movie recommendation machine prototype that combines real-world movie recommendation needs [18]. In this work, we investigate a collaborative filtering approach for binary facts known as the randomized reaction technique, which preserves privacy. We create a technique that focuses on the second privacy concern in order to determine false binary rankings by using public and auxiliary data [19]. One helpful technology that can help with the issue of users receiving too much information is recommendation systems. It makes it possible to recommend items related to the user, generates a list of recommendation rankings for each user, and predicts the grade of items to be recommended to the user [20]. By implementing a recommendation system, a number of platform services actively suggest tailored products that satisfy users' needs. Studies on different recommendation filtering models and data mining techniques are being carried out in an effort to enhance the effectiveness of these recommendations [21]. Hasan (2022) claims that recommendation models are algorithms designed to make it easier for users to find precise and pertinent products by sifting through a sizable information database. These frameworks evaluate customer choices to identify patterns in the data set and generate results that align with their individual requirements and interests. According to Nasri (2022), the main objective of recommendation systems is to predict user preferences and suggest films that they are likely to find interesting. Recommendation systems can predict whether a specific user would prefer an item based on their profile and interests, though they are primarily used in commercial settings. Kumar et al. [22] developed MOVREC, a movie recommendation system based on collaborative filtering techniques. This type of filtering utilizes data from an entire user base in order to produce recommendations. De Campos et al. [23] also performed an evaluation of both the conventional recommendation methods. Since both of these methods have some delays, he suggested another system that is a blend of Bayesian network and collaborative method. Clustering was suggested by Kużelewska [24] as a way to address the suggestions. We looked at memory-based clustering methods and centroid-based solutions. Consequently, specific recommendations were generated. Sharma and Maan [25] examined a range of recommendation methods in their paper, such as collaborative, content-based, and hybrid recommendations. It also describes the advantages and drawbacks of these techniques. Li and Yamada proposed an algorithm for inductive learning [26]. In this case, a tree that shows user suggestions has been constructed. A few of the most significant additions to the recommendation system are examined in Table 1. http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 470 Table 1: Recommendation Systems literature Review Authors Year Descriptions Scharf & Alley [27] 1993 To forecast the ideal fertilizer rate for winter wheat, the authors put forth a versatile multicomponent rate recommendation system. Basu et al. [28] 1998 The authors propose recommendation strategies that utilize both content facts and ratings. Sarwar and associates. [29] 2001 Several methods for determining item-to-item resemblance were mentioned by the authors. Bomhardt [30] 2004 The author suggested a method for making individualized news recommendations. Manikrao & Prabhakar [31] 2005 A dynamic web selection framework was developed and offered by the authors. Von Reischach et al. [32] 2009 The rating concept put out by the authors permits users to create their own rating criteria. Choi and associates. [33] 2012 The authors offered methods for integrating various tactics to improve the caliber of proposals. Table 2. covered the role of filtering Methods for various objectives. Authors Year Descriptions Goldberg and associates. [34] 1992 The team-based analyzing technique was offered by the authors. Herlocker et al. [35] 1997 The writers employed filtering strategies on Usenet news. Miyahara & Pazzani [36] 2000 The authors described a technique for automatically determining how similar a user's positive and negative ratings are. Hofmann [37] 2004 An entirely new family of model-based algorithms was presented by the author. Dabov et al. [38] 2008 The authors recommended a collaborative filtering method for image restoration. Pennock et al. [39] 2013 The authors offered a number of techniques for personality diagnosis filtering. Liu et al. [40] 2014 The authors stated an innovative way to offer a precise thought. http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 471 Methodology: This study used a hybrid methodology that combines cosine similarity and text-tonumber conversion. A recommendation technique called cosine similarity methodology forecasts and suggests highly regarded films to users by contrasting and comparing user behavior and preferences. The Alternating Least Square algorithm was used to improve the methodology. The latent attributes that capture the underlying relationships and patterns in the dataset can be extracted more easily thanks to this factorization algorithm. The matrix factorization method helps identify the latent factors that influence user preferences and movie ratings by using Alternating Least Squares. This enables the system to produce tailored suggestions for users and make precise predictions. The Suggested Model: The suggested model is a hybrid movie recommendation system that makes use of cosine similarity and text-to-number conversion. Hybrid recommendation frameworks combine several strategies to increase the accuracy of movie recommendations. This study's model-based methodology made use of matrix factorization and the Alternating Least Squares algorithm (ALS). The broad and diverse data set that will comprise the proposed framework will comprise a substantial collection of films. This dataset will include representations of language, genres, release years, and other relevant attributes. It was thoughtfully selected and arranged to guarantee that a variety of films from different sources were included. The dataset will be carefully assembled, incorporating films from a variety of genres, including documentaries, sci-fi, romance, action, adventure, comedy, thriller, and more. In order to provide comprehensive coverage of the film industry, it will cover both well-known and obscure films. Additionally, the dataset will include films in a variety of languages, allowing recommendations for users with varying linguistic preferences and capturing a global audience. The framework will be able to provide users from various cultural backgrounds with tailored movie recommendations thanks to this multilingual component. Figure 1. Work flow of the Planned Movie Recommender System Figure 1 illustrates a typical movie-recommendation pipeline: start with a raw Dataset, which goes through Preprocessing (including Exploratory Data Analysis, http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 472 Data Cleaning, and Feature Extraction) to produce usable features. These features then feed into a modeling stage where a Hybrid Model is trained and validated (the training and validation loop), suggesting a combination of methods to improve accuracy. The process culminates in a Movie Recommendation output. Dataset: The suggested model will use the popular and extensively used "TMDB 5000 Movie Dataset" from the Kaggle dataset for the experiments. An established database that is frequently used in the analysis and assessment of recommendation systems is the TMDB 5000 Movie Dataset. Because it incorporates user ratings and community user data, it is a useful tool for investigating user preferences and generating recommendations. The TMDB 5000 Movie Dataset is a comprehensive collection of user ratings that provides information about how people watch and rate movies. The vast number of users who have rated films across a wide range of genres allows for a detailed analysis of user preferences and behavior. The Movie dataset used in this proposed model is a comprehensive collection of 4,800 films. Because users in the dataset were chosen at random, a diverse range of movie preferences was represented in the user base. To verify the model's adequacy and reliability, the recommendation system will be applied to the 4,800 films. One movie title should suggest five related films. This minimal criterion guarantees that users have provided enough movie suggestions for an accurate assessment of their preferences. Individual user analysis and customized recommendation generation are made possible by the distinct user IDs assigned to each dataset participant. Experiment and Results: Table 1. Logistic Regression Class Precision Recall F1-Score Support Negative 0.86 0.83 0.84 2475 Positive 0.84 0.86 0.85 2525 Accuracy 0.8476 The table 1 presents the performance metrics of a Logistic Regression model used for binary classification, detailing its effectiveness in distinguishing between negative and positive classes. Precision measures the accuracy of positive predictions, where the negative class has a precision of 0.86 while the positive class has 0.84. Recall shows how well the model can detect real positive examples; in this case, the positive class's recall was higher at 0.86 than the negative class's at 0.83. The F1 score, which strikes a balance between recall and precision, is 0.85 for the positive class and 0.84 for the negative class. With an overall accuracy of 0.8476, the model demonstrates a substantial level of correctness across the dataset, which consists of 2,475 instances for the negative class and 2,525 for the positive class. http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 473 Table 2. Decision Tree Class Precision Recall F1-Score Support Negative 0.70 0.71 0.71 2475 Positive 0.71 0.70 0.71 2525 Accuracy 0.7076 The table 2 presents the performance metrics of a Decision Tree classifier for predicting two classes: Negative and Positive. For the Negative class, based on a support of 2,475 instances, the precision, recall, and F1_score are 0.70, 0.71, and 0.71, respectively. With an F1_score of 0.71 from a support of 2,525 instances, the metrics for the Positive class also show somewhat higher precision (0.71) and recall (0.70%). The model's overall accuracy, or the percentage of true results relative to the total number of cases examined, is 0.7076. This suggests that the Decision Tree maintains relatively balanced performance across both classes. Table 3. Random Forest Class Precision Recall F1-Score Support Negative 0.82 0.84 0.83 2475 Positive 0.84 0.82 0.83 2525 Accuracy 0.8310 Performance metrics for a Random Forest model assessing the Negative and Positive classes are shown in table 3.It contains important metrics like F1_score, Precision, and Recall in addition to Support, which shows how many instances of each class were classified. The model obtained an F1_score of 0.83 with 2,475 instances for the Negative class, with a Precision of 0.82 and a Recall of 0.84. The Positive class, on the other hand, had a lower Recall of 0.82 and a higher Precision of 0.84, resulting in an F1_score of 0.83 with 2,525 instances. With an overall model accuracy of 0.8310, both classes are performing well. Bookmark message Copy message Export. Table 4. SVM Class Precision Recall F1-Score Support Negative 0.88 0.85 0.86 2475 Positive 0.85 0.89 0.87 2525 Accuracy 0.8668 The performance metrics of a Support Vector Machine (SVM) model for binary classification into positive and negative classes are shown in table 4. Precision, which displays 0.88 for negative and 0.85 for positive predictions, gauges how accurate the positive predictions are. Recall, which is 0.85 for negative classes and 0.89 for positive classes, shows how well the model can identify all pertinent instances. A good balance between precision and recall is indicated by the F1 score, which is 0.86 for negatives and 0.87 for positives. The model's overall accuracy is 0.8668, meaning that roughly 86.68% of its predictions come true. The support values represent the number of true instances for each class, with 2475 negatives and 2525 positives, providing insight into the distribution of the classes in the dataset. http://amresearchreview.com/index.php/Journal/about Volume 3, Issue 11 (2025) Online ISSN Print ISSN . . http://amresearchreview.com/index.php/Journal/about Page 474 Table 5. KNN Class Precision Recall F1-Score Support Negative 0.54 0.39 0.45 2475 Positive 0.53 0.68 0.60 2525 Accuracy 0.5354 The table 5 presents evaluation metrics for a K-Nearest Neighbors (KNN) classification model, detailing its performance on two classes: Negative and Positive. With a precision of 54% for negative cases and 53% for positive cases, the model demonstrates the accuracy of the positive predictions. Recall gauges the model's capacity to recognize real positive examples; it performs worse for negative cases (39%), but better for positive cases (68%). The F1 score, which balances precision and recall, is 0.45 for Negative and 0.60 for Positive, indicating moderate performance overall. The model's overall accuracy stands at approximately 53.54%, which reflects how well it performs across both classes. The number of actual instances in each class is indicated by the Support values, which are 2,475 for Negative and 2,525 for Positive. Table 6. XGBoost Class Precision Recall F1-Score Support Negative 0.86 0.80 0.83 2475 Positive 0.81 0.87 0.84 2525 Accuracy 0.8350 The performance metrics for a predictive model that uses XGBoost to evaluate two classes Negative and Positive are shown in table 6. The model obtained an F1 score of 0.83, a precision of 0.86, a recall of 0.80, and a support of 2475 instances for the Negative class. On the other hand, the Positive class had a support of 2525 instances and a precision of 0.81, recall of 0.87, and F1 score of 0.84. With a reported overall model accuracy of 0.8350, the model performs well in differentiating between the two classes. These metrics show how well the model balances precision and recall for both classes while producing accurate predictions.