scieee AI-readable full text Open interactive document viewer

A QUANTITATIVE CORRELATIONAL STUDY OF ASSESSING THE EFFECTIVENESS OF RANDOM FOREST IN REDUCING FALSE POSITIVES IN INTRUSION DETECTION SYSTEMS FOR ENTERPRISE NETWORKS

Candelario, Jhan Kyle M.; Ruiz, Kirstien Sunday T.; Bulante, Charles Derick P.; Medina, Tyrone Justine A.; Mag-isa, Joan C.

Abstract

The Intrusion Detection Systems (IDS) have been vital in ensuring that enterprise networks are not affected by cyber threats by checking traffic and detecting abnormal activities. Although important, IDS is likely to have high false positive rates that may bombard security staff and decrease efficiency. Machine learning has come up as a promising solution to this problem, where the Random Forest (RF) has been identified as an ensemble-based technique where the detection is enhanced and the number of false alarms reduced. This paper analyses how TUP-T students perceive the effectiveness of the Random Forest in reducing false positive in IDS in enterprise networks. The qualitative research design was employed under descriptive research design in which a survey questionnaire was viewed as the prime data collection tool. The survey was also distributed to 100 participants through the Google Forms over the 2 weeks and contained both demographics and the knowledge of the application of the IDS and the Random Forest, attitudes towards the usefulness of the RF and the perceived benefits and problems of its utilization as well. The quantification of respondent measures was performed on Likert scale and quantification of data was done using the assistance of descriptive statistics via frequencies, percentages/ mean scores. These types of ethical concerns such as voluntary participation, anonymity and confidentiality were thoroughly observed. The results revealed a positive mark on the level of awareness of the interviewees since most of them expressed a high level of agreement of the claims that random forest improves precision of IDS, reduces false positives, and even has the ability to perform superiorly as compared to traditional instruments of detection. Other benefits that were identified by the participants included improved efficiency, less workload, and greater confidence in the safety of enterprise security mechanisms, as well as the resource needs and the complexity of the implementation process. The research adds a human aspect to the current literature on the technical research with respect to showing that user perceptions are corroborated with empirical research on the effectiveness of Random Forest. In general, the results indicate that Random Forest has a great potential to be applied to enterprise IDS, which enhances the trust and credibility of cybersecurity.

Full text

cognizancejournal.com Candelario, Jhan Kyle M. et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 411-417 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 411 A QUANTITATIVE CORRELATIONAL STUDY OF ASSESSING THE EFFECTIVENESS OF RANDOM FOREST IN REDUCING FALSE POSITIVES IN INTRUSION DETECTION SYSTEMS FOR ENTERPRISE NETWORKS Candelario, Jhan Kyle M.; Ruiz, Kirstien Sunday T.; Bulante, Charles Derick P.; Medina, Tyrone Justine A.; Mag-isa, Joan C. TUPT-NS-T-4A-T, Group 1, Technological University of the Philippines, Taguig, Km 14 East Service Road, Western Bicutan, Taguig City, Philippines DOI: 10.47760/cognizance.2025.v05i10.038 Abstract: The Intrusion Detection Systems (IDS) have been vital in ensuring that enterprise networks are not affected by cyber threats by checking traffic and detecting abnormal activities. Although important, IDS is likely to have high false positive rates that may bombard security staff and decrease efficiency. Machine learning has come up as a promising solution to this problem, where the Random Forest (RF) has been identified as an ensemble-based technique where the detection is enhanced and the number of false alarms reduced. This paper analyses how TUP-T students perceive the effectiveness of the Random Forest in reducing false positive in IDS in enterprise networks. The qualitative research design was employed under descriptive research design in which a survey questionnaire was viewed as the prime data collection tool. The survey was also distributed to 100 participants through the Google Forms over the 2 weeks and contained both demographics and the knowledge of the application of the IDS and the Random Forest, attitudes towards the usefulness of the RF and the perceived benefits and problems of its utilization as well. The quantification of respondent measures was performed on Likert scale and quantification of data was done using the assistance of descriptive statistics via frequencies, percentages/ mean scores. These types of ethical concerns such as voluntary participation, anonymity and confidentiality were thoroughly observed. The results revealed a positive mark on the level of awareness of the interviewees since most of them expressed a high level of agreement of the claims that random forest improves precision of IDS, reduces false positives, and even has the ability to perform superiorly as compared to traditional instruments of detection. Other benefits that were identified by the participants included improved efficiency, less workload, and greater confidence in the safety of enterprise security mechanisms, as well as the resource needs and the complexity of the implementation process. The research adds a human aspect to the current literature on the technical research with respect to showing that user perceptions are corroborated with empirical research on the effectiveness of Random Forest. In general, the results indicate that Random Forest has a great potential to be applied to enterprise IDS, which enhances the trust and credibility of cybersecurity. Keywords: Intrusion Detection System (IDS), False Positives, Random Forest, Machine Learning, Enterprise Networks, Perception Study. cognizancejournal.com Candelario, Jhan Kyle M. et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 411-417 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 412 I. INTRODUCTION The IDS are considered to be one of the most crucial systems to secure enterprise networks since they constantly scan the network traffic and system activity to seek any form of anomalies, violation of security policies or even the attackers. [1]. However, after a critical role, there are high false alarms in Intrusion Detection Systems that flood the security staff with irrelevant alarms and thus the creativity to develop more realistic detection models shall remove the number of false alarms and sustain high detection accuracy [2] [3]. The machine learning (ML) algorithms have become the future to revolutionize the functionality of the IDS especially to address the noisy, heterogeneous, and imbalanced nature of the network data. [4]. One of this is the Rand Forest (RF) algorithm, though the key mark of this algorithm is that an algorithm is an assembly; in other words they are a collection of decision-trees whose consequence is that the classification law would offer better results throughout the classification process. [5]. RF has also been proven to perform better than their well-known algorithms in the detection of anomaly. [6]. The paper is a quantitative and experimental comparison of implementation of both the Random Forest and IDS on benchmarks dataset, NSL-KDD and UNSW-NB15. All measures of performance. These findings show that the false positive detection would be lowered considerably when using the Random Forest and consequently allow to increase the reliability and relevance of the IDS alerts to individual security operators of an enterprise. [8]. Along with that, the Random Forest algorithm is likewise appropriate due to its flexibility and scalability that need to be further implemented in a real-time monitoring of enterprises networks so that additional investigation of peripheral integrative features to establish consistency with deep learning models as well as deploying the algorithm on Intrusion Prevention System. [9]. The resiliency of IDS and the better defense of enterprise-level cybersecurity in dynamically demanding threat environments should be demonstrated to be better through such explainable hybrid models than before. [10]. II. METHODOLOGY A. Research Methodology in Intrusion Detection System (IDS) Studies The approaches employed in this study rely on the methodological models applied in past studies in the area of IDS. The work by Tang et al. [11] is a systematic analysis of machine learning algorithms in enterprise IDSs. To compare their performance level in relation to a false positive in determining the level of effectiveness of the adaptive algorithms in real-life network situation, they designed their experiment by naively using adaptive algorithms in realistic network environments and quantifying differences statistically in order to determine which discrepancies are of any significance. Similarly, the article of Abdallah et al. The survey Article survey was carried out on the Random Forest-based IDS techniques which is concerned with predictive power and explainability [12]. The authors benchmarked the NSL-KDD and UNSWNB15 trainer and the results were measured using False Positive Rate and True Positive rate (TPR) as an evaluation of the algorithm performance. Besides these, one study by Benaddi et al. [13] proposed an adaptive machine learning method founded upon the Random Forest and ensemble methods in finding and mitigating false alarms in IDS. Off-line training and realtime testing phase were affected by the use of the methodology to determine statistically significant performance change. The character of correlation studies they are involved in regarding the parameters of the algorithm and the outcomes of detecting them represent the applicability of correlational designs in the specified research. Average time, the meta-analytic synthesis adopted by Liu et al. [14] to analyze empirical research data on IDS provided insight into the usefulness of the meta-analytic synthesis research method by synthesizing 54 empirical studies to generate generalized findings on the trust, risk, and machine learning performance. The reason why their moderator analyses were prominent is that they made it a point to highlight the considerations of the contextual factors such as network environment, kind of a dataset. B. Sequential (Waterfall) Methodology Framework The study methodology can be considered sequential (Waterfall), as the major parts that needed to be organized to conduct the study of the relationship between parameters of the Random Forests and the rates of false positives were the enterprise network IDS as the main parameters. cognizancejournal.com Candelario, Jhan Kyle M. et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 411-417 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 413 Phase 1: Study Design and Data selection. The study submitted is of the quantitative and correlational character, and it is focused at identifying the correlation between the work of a Random Forest and the enterprise network environment. Adoption of NSLKDD and UNSWNB15 benchmark Intrusion Detection System (IDS) data set data is appropriate, as the two are common and have been extensively applied in carrying out intrusion detection studies. The writing cases are the ones that have all the feature set and are definitely identified and the unexplained or the corrupted data cases are disposed of. There was also a given sample size of 100,000 or more instances of traffic to get the necessary statistical power that would be utilized during a correlation and regression analysis [11]. Phase 2: Design and Engineering Algorithms and Feature. The main intrusion detecting algorithm is the random forest that is developed. The feature selection based on the standard methods [12][13] is linked with the nature of a network traffic such as the type of protocols, number of packets, number of flags, duration and bytes. The meaning of features is extracted in the form of gini impurity rating of the most meaningful features which are the ones rated to have the highest scores. Pre-processing involves processes such as information undergoing standardization, normalization and conversions. Phase 3 Data Handling and system deployment Ethics. The study says that the paramount concepts of privacy of data and system security are practiced in an amazing way, yet the datasets are open and available. There are such limits to its data usage that are liable to its licenses and within a limited virtual enterprise configuration. There is no loss of the experimental environment as only the local servers where all the data are processed are actually encrypted. It is simulated and employs real-time deployment of IDS i.e., the traffic cases are processed via the trained Random Forest real life model wherein inputs are adjustable in a confined lab network, therefore, this qualifies as a simulation of the ISO 20252:2019 protocols of Data handling [16]. Phase 4: Transformation and Cleaning of Data. There is a systematic and orderly organization of cleaning and transformation of raw network data. Duplications and entries that lack completeness are removed together with the anomalies in the data. Continuous variables are variegated and categorical variables are one-hot coded. False Positive (FP) and True Positive (TP) similar outcomes can be seen in binary codes (0 represents that it has not been detected, 1 accords that it has been detected). False positive rate is performed session-wise, dataset characteristics are normalized to be statistically tested. Phase 5: Statistical Analysis and Implementation. To examine the relationship between the values of the parameters of Random Forests and the rate of the use of false positives, it moves with the help of a number of quantitative methods. The description statistics gives a summary of all the peculiarities of the data of the dataset and distributions of the discerningness of the models. Point-biserial correlation is used to compare correlation between an important hyperparameter (number of estimators"), versus binary classification data. False positives are likely to occur through the most significant predictors that have been identified considering binary logistic regression. The differences in the unveiling of several parameters settings at = 0.05 are being insinuated by the chi-square tests. The methods used to identify the trends of the traffic types yielded a false positive of the type of traffic using the K-means clustering algorithm using the studies by Tang et al. [11] or Benaddi et al. [13] as specifications. Phase 6: Interpretation and Dissemination. The final outcomes of the discussion are the outcomes of performance analysis of machine learning in the setting of the Intrusion Detection System (IDS) and they significantly predetermine the topicality of the phenomenon of alleviation of the instances of false. positives in the practice of security in an enterprise. cognizancejournal.com Candelario, Jhan Kyle M. et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 411-417 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 414 Every conclusion is related to the questions of the study and the support of the form of the tables, graphs, and statistical suspicions is provided. The limitations segment is where the area of the representativeness of the dataset and the usage of the simulated environment are observed. It informs the implications on the network administrators, direction of research in future, as well as security policy planners such as, hybrid ensemble schemes, and adaptive tuning policies. III. RESULT AND DISCUSSION A. Respondent Profile and Awareness Figure 1. Awareness of IDS and Random Forest The survey study aimed at targeting 100 students of TUP-T that were affiliated with the discipline of information technology and cybersecurity. The population mix in their favor was that they were assured of the sample population which in this case would comprise of people who already had information on the Intrusion Detection Systems (IDS) and machine learning concepts. Most of the respondents were experienced in IDS and the challenges in its operations. Extraordinarily, 95 percent of them got worried about this fact that high false positive rates savored efficiency within the security department and it is affected by 94 percent that machine learning existed within security systems. The very pretty 98 percent responded in the affirmative that they agreed this is true with the general awareness rates being high and very importantly 98 percent stated that they strongly agreed that the model of random forest is a good model to detect cyber threats. B. Perceived Effectiveness of Random Forest Most of the respondents supported the application of the Random Forest to optimize the IDS by 96 percent, 96 percent, 99 percent, respectively, and thought that offering prediction capabilities, dead air, and even better to call resource utilization and time-saving respectively. These impressions can be aligned to the literature sources available which persist in reinforcing the fact that it is not only possible to obtain high values of detection using Random Forest but it is also possible to begin to decrease the figures of false alarms using benchmark data sets that are namely NSL-KDD and UNSWNB15. The high level of reliability of various items proves that the respondents view the concept of Random Forest as one of the solutions to the current issue of false positives in enterprise networks. cognizancejournal.com Candelario, Jhan Kyle M. et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 411-417 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 415 Figure 2. Perception on Effectiveness C. Benefits, Challenges, and Interpretation The respondents stated the primary benefit of the false positives minimization which, consequently, also leads to the growth of efficiency, shortened response time, and the enhancement of the trust in the safety of the enterprises. In fact, almost all the respondents (98% thought that the key to building confidence in the work of IDS is reducing the number of false alarms, and 96% acknowledged that AI can help a lot to improve the detection rates). These findings indicate that such awareness is a major variable in perception according to which the people who know well view Random Forest as a more feasible option in place of IDS and machine learning. Therefore, it suggests that the potential employees can be altered in their attitude toward machine learning-based IDS solutions through the assistance of pre-conferenced educational procedures such as awareness campaigns, training, or curriculum. Figure 3. Benefits and Challenges Identified However, the issues of the process of implementation elaboration and resource consumption were also mentioned among the respondents. Even though the overall attitude of the people toward a general implementation of the concept of the Random Forest in the shelves of the organization is currently quite positive, these concerns bring about the reality that, in addition to having a proper infrastructure that will enable such connectivity, it should also have the proper skilled human resource in place that will enable the process. D. Implications and Literature Comparison It is a good omen that the present perceptions of the respondents are favorable because it is likely that in the future, the IT professionals will be accommodating when it comes to the application of the Random Forestbased IDS on the enterprise level. The fact that the IT community has recognized such a stance is a motivational factor of adaptation of technology and thus, organizations can be motivated to apply the tactics of machine cognizancejournal.com Candelario, Jhan Kyle M. et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 411-417 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 416 learning to reduce the rate of false positives and have the quality of detection rates that significantly outperforms the rate obtained using the ancient IDS methods. The other good thing that makes the argument on the theory and the practical efficiency of the algorithm true is this correlation between the perception and human outcomes of the experiments. IV. CONCLUSION AND RECOMMENDATIONS The aim of the experiment was to make a comparison of the perceptions of the students of TUPT concerning the relevance of the use of the random forest in minimizing false positives in the Intrusion Detection System (IDS). These results have proven that the respondents were very familiar with the concept of IDS and machine learning and most of them were aware of the use of the Random Forest as the method to enhance the performance as the one that surpasses the classical IDS models. These findings indicate that future IT professionals embodied by the students are positive about Random Forest, on its capacity to reduce one of the most raging issues in enterprise network securityfalse positives. Although such optimism exists, there are still other barriers to adoption that exist and they include requirement of resources, technical integration and even sufficient training. The research will be valuable as it will present a humanistic approach to the technical literature regarding the use of Random Forest that demonstrates that the perceptions follow the empirically tested evidence of its utility. In general, the optimism expressed by the respondents discusses the further development of the Random Forest and its possible integration into the system of enterprise IDS that would lead to more credible and trustworthy cyber security systems. ACKNOWLEDGEMENT The researchers would like to say they highly appreciate the individuals who assisted in the success of this study. Above all is the gravitas given to that fact that the students of TUPT who were not paid to take part in the survey but had committed their time, wisdom, and opinion to do so as volunteers and at their own free will. The basis of this study in determining the effectiveness of the Random Forest in minimizing the impact of false positives on Intrusion Detection Systems on the enterprise networks lay on their responsiveness. The respondents were very cooperative and beautiful and this is why the research was a great success. The researchers also introduce the fact that they cannot show their gratitude to their academic advisor, people who were on the panel, mentors and people that were extremely helpful and constructive and patient to complete this research. Similarly, family, friends, and colleagues who in one way or other was always around, supportive, understanding are also given special attention and appreciation as such. Above all, the researchers owe an immeasurable debt of gratitude to their entire workforce of people who in one way or other contributed to the success of the study achievement to take placewhether in large or small ways. REFERENCES 1. Abdelaziz, M. T., Mohamed, A., El-Sayed, A., & Hassan, H. (2025). Enhancing network threat detection with Random Forest. Journal of Communications and Networks . https://doi.org/10.1007/s10922-024-09874-0 2. Sowmya, T., et al. (2023). A comprehensive review of AI based intrusion detection. Retrieved from ScienceDirect. 3. Arreche, O., Bibers, I., & Abdallah, M. (2024). A comprehensive comparative study of individual machine learning models and ensemble strategies for network intrusion detection systems. arXiv preprint arXiv:2410.15597 . https://arxiv.org/abs/2410.15597 4. Chavan, R. (2025). A high-accuracy approach to intrusion detection using CIC-IDS 2017. International Journal of Computer Applications, 187 (3), 1–6. https://ijcaonline.org/archives/volume187/number3/chavan-2025ijca-924816.pdf 5. Doost, P. A., Sarhani Moghadam, S., Khezri, E., Basem, A., & Trik, M. (2025). A new intrusion detection method using ensemble classification and feature selection. Scientific Reports, 15 , 98604. https://doi.org/10.1038/s41598-025-98604-w 6. Nassreddine, G., Nassereddine, M., & Al-Khatib, O. (2025). Ensemble learning for network intrusion detection based on correlation and embedded feature selection techniques. Computers, 14 (3), 82. https://doi.org/10.3390/computers14030082 7. Patrick, B. (2025). Reducing false positives in intrusion detection systems with adaptive machine learning algorithms. ResearchGate . https://www.researchgate.net/publication/390747122 cognizancejournal.com Candelario, Jhan Kyle M. et al, Cognizance Journal of Multidisciplinary Studies, Vol.5, Issue.10, October 2025, pg. 411-417 (An Open Accessible, Multidisciplinary, Fully Refereed and Peer Reviewed Journal) ISSN: 0976-7797 Impact Factor: 5.183 Index Copernicus Value (ICV) = 92.57 ©2025, Cognizance Journal, ZAIN Publications, Fridhemsgatan 62, 112 46, Stockholm, Sweden, All Rights Reserved 417 8. Rai, H. M. (2024). Improved network intrusion detection techniques using feature optimization. Mathematics, 12 (24), 3909. https://doi.org/10.3390/math12243909 9. Larriva-Novo, X., et al. (2023). Leveraging explainable artificial intelligence in real-time intrusion detection systems. MDPI Applied Sciences. Retrieved from https://www.mdpi.com/20763417/13/15/8587 10. “Enhancing intrusion detection systems with ensemble learning techniques in machine learning.” (2024). ResearchGate . https://www.researchgate.net/publication/384473626 11. Tang, C., Wang, S., Wang, L., & Song, Y. (2024). Reducing False Positives in Intrusion Detection Systems with Adaptive Machine Learning Algorithms. ResearchGate. https://www.researchgate.net/publication/390747122_Reducing_Fals e_Positives_in_Intrusion_Detection_Systems_with_Adaptive_Machi ne_Learning_Algorithms 12. Abdallah, A., Maarof, M. A., & Zainal, A. (2018). A Survey of Random Forest-Based Methods for Intrusion Detection Systems. ResearchGate. https://www.researchgate.net/publication/323129609_A_Survey_of_ Random_Forest_Based_Methods_for_Intrusion_Detection_Systems 13. Benaddi, M., El Ouahidi, B., & Abdelhadi, A. (2023). Reducing False Positives in IDS Using Adaptive Random Forest. arXiv preprint. https://arxiv.org/pdf/1801.02330 14. Liu, F., Handoyo, S., & Park, S. (2024). Meta-analysis of Machine Learning Algorithms in IDS: Trust, Risk, and Performance. MDPI Applied Sciences, 14(2), 714. https://www.mdpi.com/20763417/14/2/714 15. Moustafa, N., & Slay, J. (2015). UNSW-NB15: A Comprehensive Data Set for Network Intrusion Detection Systems. Military Communications and Information Systems Conference (MilCIS), 1–6. 16. ISO. (2019). ISO 20252:2019 – Market, opinion and social research, including insights and data analytics — Vocabulary and service requirements. International Organization for Standardization.