scieee AI-readable full text Open interactive document viewer

AI-BASED DETECTION AND MITIGATION OF SPAM, BOTS, AND MALICIOUS WEB TRAFFIC

R.K. Pirova, Sh.K. Shoykulov

Abstract

Automated spam generation, bot-driven attacks, and malicious web traffic pose serious risks to modern online platforms, causing performance degradation, data leakage, and financial losses. This study presents an AI-driven framework for detecting and mitigating harmful traffic using supervised machine learning, behavioral analytics, and anomaly detection methods. The proposed approach combines feature-rich traffic profiling with hybrid classification models to distinguish legitimate sessions from automated or adversarial ones. Experimental results on real and synthetic datasets demonstrate high detection accuracy, reduced false-positive rates, and improved interpretability through Python-based visual analytics. The findings confirm that AI provides a scalable and effective defense mechanism for securing web infrastructures against evolving cyber threats.

Full text

SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 11 NOVEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 103 AI-BASED DETECTION AND MITIGATION OF SPAM, BOTS, AND MALICIOUS WEB TRAFFIC R.K. Pirova1, Sh.K. Shoykulov2 PhD1 Associate Professor2 Department of Applied Mathematics, Karshi State university, Republic of Uzbekistan1,2 https://doi.org/10.5281/zenodo.17799296 Abstract. Automated spam generation, bot-driven attacks, and malicious web traffic pose serious risks to modern online platforms, causing performance degradation, data leakage, and financial losses. This study presents an AI-driven framework for detecting and mitigating harmful traffic using supervised machine learning, behavioral analytics, and anomaly detection methods. The proposed approach combines feature-rich traffic profiling with hybrid classification models to distinguish legitimate sessions from automated or adversarial ones. Experimental results on real and synthetic datasets demonstrate high detection accuracy, reduced false-positive rates, and improved interpretability through Python-based visual analytics. The findings confirm that AI provides a scalable and effective defense mechanism for securing web infrastructures against evolving cyber threats. Keywords: AI security, web traffic analysis, bot detection, spam filtering, anomaly detection, machine learning, cybersecurity automation. INTRODUCTION In recent years, the web has been confronted with a rapid increase in automated threats: spam floods, malicious bots, scripted attacks, and atypical request behavior have become an integral part of the digital environment. These phenomena lead to server overload, service disruption, data leakage risks, and direct financial losses, particularly for commercial and financial web platforms [1], [2]. Regular research confirms that the share of non-human traffic is increasing annually, and the tools used by attackers are becoming increasingly accessible and technologically advanced [3]. Classical approaches to threat filtering—static signatures, blacklists, CAPTCHAs, and simplified request rate analysis rules—are no longer able to counter adaptive attacks. Modern malicious agents mask identifiers, imitate human actions, distribute the load across multiple sources, and exploit dynamic behavior patterns, making traditional defenses increasingly ineffective [4], [5]. In such conditions, intelligent data analysis methods capable of identifying hidden patterns in traffic structure and detecting attacks without relying on rigidly defined rules are in demand. The advent of advanced machine learning and artificial intelligence algorithms has significantly expanded the capabilities of network log analysis. Classification models, anomaly detection methods, and neural networks trained on query sequences enable a more accurate understanding of web activity. Several studies demonstrate the high effectiveness of gradient boosting, deep learning models, and autoencoders in detecting malicious sessions, as well as the ability of AI to identify previously unknown types of attacks [6], [7], [8]. However, questions SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 11 NOVEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 104 regarding the interpretation of such models, their scalability, and resilience to changing attacker strategies remain open. This study aims to develop a comprehensive approach that combines supervised learning methods, user behavior analysis, and anomaly detection algorithms. The scientific novelty lies in the creation of a multilayer architecture where classifiers, sequence models, and unsupervised algorithms operate synchronously, assessing the risk of each individual session. This approach ensures resilience to a variety of malicious traffic types and reduces the likelihood of false positives. The goal of the work is to build an AI-based technology for the automatic detection and filtering of malicious web traffic. To achieve this, the following tasks are addressed: generating a representative dataset, extracting informative features from HTTP sessions, training hybrid models, comparing their effectiveness, and visually presenting the results [9], [10]. The object of the study is web traffic on modern online platforms, and the subject is AI methods used for classification, ranking, and detection of abnormal behavior. RESULTS and DISCUSSIONS The study was based on a comprehensive application of data analysis methods, artificial intelligence tools, and visual analytics. The analysis was based on a combination of real HTTP traffic and artificially constructed malicious sessions. The real request logs contained timestamps, HTTP protocol parameters, information on transferred data volumes, latencies, client types, and access paths. All sensitive elements were anonymized before processing. The synthetic traffic was generated to reproduce the most common attack patterns. It included high-frequency sequences characteristic of botnets, automated attempts to brute-force credentials, the generation of similar spam requests, and distributed attack patterns. The use of two sources ensured the flexibility of the experiment and the reproducibility of various threat scenarios. To ensure the correct operation of the machine learning algorithms, the traffic was transformed into a structured feature space. This included parameters characterizing both individual requests and client behavior within a single session. The statistical component included metrics for intensity, variability of intervals between requests, and the entropy of visited URLs. Protocol features reflected the correctness of headers, the presence of anomalous clients, and suspicious referral sources. Behavioral characteristics allowed us to identify repetitive action sequences, navigation cycles, and sudden spikes in activity. Numerical variables were normalized to a single scale, and categorical values were encoded using frequency representation methods.[6] The system architecture included several models, each performing its own level of analysis. XGBoost and Random Forest algorithms were used to classify sessions based on a set of features. An LSTM neural network, capable of capturing the dynamics of request sequences, was used to recognize temporal patterns. The anomaly detection component consisted of an autoencoder and the Isolation Forest algorithm. The former determined the degree to which the current session deviated from normal behavior based on feature recovery error, while the latter assessed the rarity of specific combinations in the data space. Combining their findings reduced the risk of misclassifying legitimate activity as malicious. The analytical process was based on sequential processing of data by multiple models. Classifiers identified the type of activity based on pre-labeled examples, while unsupervised algorithms analyzed deviations from the normal behavior pattern. This approach proved particularly effective in situations where malicious agents attempted to imitate legitimate behavior. Classification errors were further analyzed using visual reports, allowing for the identification of SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 11 NOVEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 105 edge cases and refinement of the feature set. To assess the quality of the proposed architecture, commonly used metrics were used: precision, recall, F1 score, ROC-AUC, and confusion matrix. The system's ability to detect new types of anomalies was separately assessed by analyzing the dynamics of reconstruction errors and the density distribution of sparse points [11]. The experiments were conducted in a Python environment using the TensorFlow, Scikitlearn, Pandas, and NumPy libraries. Visualization of the results was performed using Matplotlib and Seaborn, providing a detailed analysis of traffic behavior. An example of a visualization of request intensity is presented below. It helps identify spikes in activity, characteristic of automated attacks. import matplotlib.pyplot as plt import pandas as pd # Load traffic data from CSV data = pd.read_csv("traffic.csv") # Generate a line plot showing the change in request intensity over time plt.figure(figsize=(10, 5)) plt.plot(data["timestamp"], data["requests_per_min"], linewidth=1.5) plt.title("Dynamics of Request Intensity Over Time") plt.xlabel("Time") plt.ylabel("Requests per Minute") plt.tight_layout() plt.show() Fig 1. An example of visualizing query intensity Using such graphs allows for the rapid identification of anomalous zones and the confirmation of model conclusions with clear traffic patterns. Testing the XGBoost and Random Forest models showed that both architectures are capable of reliably identifying malicious sessions, but the best results were achieved by the XGBoost algorithm. Its high sensitivity and robustness to heterogeneous features ensured accuracy superior to other methods. Analysis of the confusion matrix revealed that rare false positives are more often associated with either excessively intensive but legitimate user requests (e.g., service integrations) or bots imitating real human behavior. However, the overall number of such cases remains insignificant[12]. Some experiments focused on identifying anomalous requests using an autoencoder and Isolation Forest. Both methods proved useful for capturing atypical combinations of features, but their combined use provided the most consistent results. The combined effect of the two algorithms allowed for the highly accurate identification of unusual patterns that are not always present in the SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 11 NOVEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 106 labeled sample. Graphical presentation of results is an important complement to numerical metrics. Line graphs reflecting the dynamics of request intensity allow for quick identification of activity periods that coincide with detected anomalies. The following version of the script was used to display anomalies and normal traffic: anomalies = data[data["anomaly_flag"] == 1] normal = data[data["anomaly_flag"] == 0] plt.figure(figsize=(10, 5)) plt.plot( normal["timestamp"], normal["requests_per_min"], linewidth=1, color="black", label="Normal traffic" ) plt.scatter( anomalies["timestamp"], anomalies["requests_per_min"], marker="x", s=25, color="black", label="Anomalies" ) plt.title("Request Intensity with Detected Anomalies") plt.xlabel("Time") plt.ylabel("Requests per Minute") plt.legend() plt.tight_layout() plt.show() Fig 2. Displaying anomalies and normal traffic This visualization demonstrates the difference between normal and suspicious behavior not only quantitatively but also graphically. A systematic analysis of all results demonstrates that the proposed hybrid architecture offers significant advantages over traditional methods of detecting malicious traffic. The model effectively classifies known attack scenarios, while the anomaly module identifies previously unobserved behavior patterns. Python graphical tools provide SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 11 NOVEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 107 additional clarity in interpreting the algorithms' actions and serve as a bridge between theoretical analysis and practical application [13]. The results obtained during the study confirm that the combined AI architecture can significantly improve the effectiveness of detecting malicious activity in web traffic. The combined use of supervised learning methods, anomaly algorithms, and behavioral analysis enabled the identification of various categories of attack scenarios, including both known and previously unseen patterns. This approach demonstrates an advantage over traditional filters, which typically respond only to predefined signatures and are ineffective when faced with changing threat types. A key observation is that XGBoost performed reliably on the classification task thanks to its ability to interpret complex dependencies between features. The obtained ROCAUC values demonstrate that the model reliably distinguishes between legitimate sessions and malicious requests. This is particularly important in real-world settings, where excessive false positives can negatively impact the user experience. Unlike simple heuristics that are sensitive to noise and traffic dynamics, modern ML models maintain accuracy even in the presence of significant variations in user behavior[14]. Analysis of the anomaly detection mechanism also revealed important patterns. The combined use of an autoencoder and Isolation Forest reduced the number of incorrect classifications and made it possible to detect atypical sessions not present in the training set. This suggests that the system has the ability to adapt to new forms of attacks that were not yet considered when generating data labels. This characteristic makes the proposed approach a promising tool in a rapidly changing digital environment. Despite the models' high performance, errors still tend to cluster around edge cases. These include legitimate users demonstrating a high frequency of requests, or bots that mimic human pauses between actions. This highlights the need to expand the feature set and integrate additional sources of information, such as network, device, or geographic origin data. Expanding the context can improve the robustness of models and enhance the system's ability to discern complex behavior patterns. Visual analytics plays a significant role. The generated graphs showed that the request dynamics and anomaly distribution are easily interpreted by cybersecurity specialists. Graphical elements help quickly identify periods of anomalous activity, correlate them with threat types, and track impacts in real time. This makes visualization an important complement to numerical metrics and improves the comprehension of analysis results. Despite the significant advantages of the proposed architecture, several limitations must be considered. Fully training the models requires a sufficiently large and diverse dataset, which may prove challenging for organizations with limited infrastructure. Furthermore, using deep neural networks in a production environment requires additional computational overhead, and regularly updating the models is key to maintaining their relevance [15]. Overall, the study demonstrates the potential of using hybrid AI methods in web traffic analysis. The system demonstrates good adaptability and can be integrated into a wide range of platforms: from online stores to banking services and government portals. The use of artificial intelligence in this area allows not only for the detection of malicious behavior but also for the proactive modeling of potential threats, thereby enhancing the security of information systems. CONCLUSION The study revealed that the use of artificial intelligence methods significantly improves the detection of malicious activity in web traffic. The developed system, which combines classification algorithms, time-sequence analysis, and anomaly detection mechanisms, demonstrates high adaptability to the variability of the digital environment and reliably identifies SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 11 NOVEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 108 both known and atypical forms of attacks. A comparison of the obtained results with traditional approaches shows that classic filters and signature-based mechanisms are significantly inferior in accuracy and flexibility. One of the main advantages of the proposed architecture is its ability to detect threats whose forms were not previously represented in the training data. This demonstrates the high potential of unsupervised learning methods and their important role in detecting previously unknown behavior patterns. The anomalous patterns detected by the system confirm that modern web services require dynamically learning defense mechanisms capable of responding to the emergence of new attacker tactics. An equally significant conclusion is the highly informative visual representations accompanying the model's operation. Line graphs and anomaly charts enable rapid assessment of traffic conditions and the identification of problematic intervals, significantly accelerating the interpretation of results by security specialists. Graphical analytics enhances model transparency and facilitates its practical application in monitoring systems. Although the proposed approach demonstrates impressive results, limitations cannot be overlooked. Training and regularly updating models require heterogeneous and relatively large datasets, and deep architectures place increased demands on computing resources. However, such shortcomings can be mitigated through architectural optimization, the implementation of autonomous retraining mechanisms, and the use of distributed computing environments. Overall, the study confirms that the integration of hybrid AI methods into web traffic analysis processes is an effective direction for the development of cybersecurity systems. The results demonstrate the potential for further developments related to expanding the feature space, improving the interpretability of algorithms, and creating more computationally efficient solutions. In the future, such systems could become a key element in protecting digital platforms in the face of the increasing complexity of cyber threats and the scale of network interactions. REFERENCES 1. Böttinger, K. (2019). Detecting bots in web traffic using machine learning. Journal of Cybersecurity and Digital Forensics, 7(3), 120–134. 2. Cao, J., Li, Y., & Liu, W. (2020). A deep learning approach to detecting malicious web requests. IEEE Access, 8, 162083–162094. 3. Chiba, Z., Abou El Kalam, A., El Ouahidi, B., & Bouhorma, M. (2017). Intelligent system for identifying malicious web traffic using supervised learning. International Journal of Network Security, 19(5), 725–736. 4. Corona, I., Giacinto, G., & Roli, F. (2013). Adversarial attacks against intrusion detection systems: Taxonomy, solutions and open issues. Information Sciences, 239, 201–225. 5. Cui, B., Jiang, S., & Yang, Q. (2021). Hybrid anomaly detection for web attacks based on autoencoder and isolation forest. Knowledge-Based Systems, 221, 106983. 6. Sh. Q. Shoyqulov, Valiyeva Sh.T. Automated technical support using ai in web environments. Евразийский журнал технологий и инноваций, p.5-12. https://doi.org/10.5281/zenodo.17656649 7. Sh. Q. Shoyqulov. On the study of optical communication systems using simulators. Eurasian journal of mathematical theory and computer sciences, Т. 5, Выпуск 11. Nov. 2025. p.20-28. https://doi.org/10.5281/zenodo.17640489 8. Sh. Q. Shoyqulov, Valiyeva Sh.T. Analysis of existing technological solutions for video surveillance systems and their limitations. Eurasian Journal of Academic Research (EJAR), Т. 5, Вып. 11. Nov. 2025. p.61-67. https://doi.org/10.5281/zenodo.17640559 SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 11 NOVEMBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 109 9. Sh. Q. Shoyqulov. AI-enhanced Web scraping for data-driven analysis. Central Asian Journal of Multidisciplinary Research and Management Studies (CAJMRMS), Vol 2, Issue 11. Nov. 2025. p.20-27. ISSN:3030-3540. www.in-academy.uz. https://doi.org/10.5281/zenodo.17529443 10. Sh. Q. Shoyqulov. Artificial intelligence for automated seo enhancement. Yangi O'zbekiston ilmiy tadqiqotlar jurnali (YOITJ), 2-jild, 11-son. IF=8.5. Nov. 2025. p.31-37. ISSN:30303559. www.in-academy.uz. https://doi.org/10.5281/zenodo.17522170 11. Sh. Q. Shoyqulov, F. S. Shodmonova. Modern algorithms for object recognition and tracking in video surveillance systems based on artificial intelligence. Science and Innovation international scientific journal, Volume 4, Issue 10, Oct. 2025. p.114-120. ISSN: 2181-3337 | SCIENTISTS.UZ. https://doi.org/10.5281/zenodo.17525448 12. Sh. Q. Shoyqulov. Integrating LLMs into Web applications: opportunities and security challenges. Eurasian journal of mathematical theory and computer sciences. Т. 5, Выпуск 6, сс. 54–60. https://doi.org/10.5281/zenodo.15755908 13. Sh. Q. Shoyqulov. AI-driven UX optimization for Web applications. Eurasian journal of mathematical theory and computer sciences. Т. 5, Выпуск 6, сс. 46–53. https://doi.org/10.5281/zenodo.15755881 14. Sh. Q. Shoyqulov, G.Kh. Astanaqulova. A comparative study of AI-based personalized recommendation algorithms for E-commerce platforms. SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4, ISSUE 5, 2025. p.142-148. ISSN: 2181-3337 | SCIENTISTS.UZ. https://doi.org/10.5281/zenodo.15589779 15. Sh. Q. Shoyqulov, E.H. Qurbonova. Intelligent analysis of user behavior in Web environments using artificial intelligence algorithms. SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4, ISSUE 5, 2025. p.176-181. ISSN: 2181-3337 | SCIENTISTS.UZ. https://doi.org/10.5281/zenodo.15590204