Full text
Corresponding author: Olamotse Roland Igbape. Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0. Resilient Dataflow Intelligence (RDI): A Proactive AI Framework for Anomaly Anticipation in Cloud Databases Olamotse Roland Igbape 1, *, Oluwatoyin Olawale Akadiri 2, Chijioke Cyriacus Ekechi 3, Richards Obada Okiemute 4, Ololade Babatunde 5 and Moyinoluwa Emmanuel Idowu 6 1 Computer Science, College of Computing, Georgia Institute of Technology, Atlanta, GA, USA. 2 Department of Information Sciences, School of Information Sciences and Engineering, Bay Atlantic University, United States. 3 Electrical and Computer Engineering, Tennessee Technological University. 4 University of Benin, Nigeria. 5 Computer Engineering, Izmir Institute of Technology, Izmir, Turkey. 6 Computer Science, Ladoke Akintola University of Technology, Nigeria. Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 Publication history: Received on 10 August 2025; revised on 20 September 2025; accepted on 22 September 2025 Article DOI: https://doi.org/10.30574/gjeta.2025.24.3.0282 Abstract Cloud databases are becoming more essential to enterprise applications, but they are still susceptible to unreliable failures, performance outbursts, and corruption of data. The conventional state-of-the-art methods of monitoring are based on reactive detection of anomalies, where the problems are recognized when they have already happened. In this paper, a new framework named Resilient Dataflow Intelligence (RDI) will be presented, which is a multi-layer AI-driven framework aimed at proactively detecting anomalies in cloud database systems. RDI combines real-time stream processing, ensemble machine learning models, and predictive analytics to identify early indicators of anomalies, which disrupt operations. Tests on three leading cloud providers, AWS RDS, Google Cloud SQL, and Microsoft Azure Database, show prediction accuracy of 93.2, a downtime of 48 percent, and inference in real time of less than 15 milliseconds. Ecommerce, healthcare, and financial services. Case studies prove the increases in database reliability, operational efficiency, and cost optimization, and the average ROI is realized within less than 10 months. The study confirms RDI as a viable and scalable instrument to change the paradigm of database monitoring and transition database monitoring tools to be more proactive rather than reactive. Keywords: Anomaly Detection; Cloud Databases; Machine Learning; Predictive Analytics; Stream Processing; Ensemble Methods; Database Reliability 1. Introduction Modern digital ecosystems are founded on cloud databases that can support finance, healthcare, and e-commerce applications at a scale never before. The global cloud database is projected to reach seventy-eight point eight billion dollars by the year 2027, an indication of the importance of the systems on the operations of the enterprises [1]. Scalability, heterogeneity, and dynamism of cloud workloads have, however, made high availability and reliability more difficult to achieve [2, 3]. Conventional methods of anomaly detection are largely reactive and only detect problems once they have occurred in the form of performance problems or service outages [4]. The result of this reactive measure is a huge amount of business, as in industry reports, the average downtime costs have been estimated to be more than 300,000 dollars per
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 315 hour in large enterprises [5]. Industries that cannot afford to lose and where the losses are mission-critical, including financial services and healthcare, suffer even greater losses, with some organizations recording costs of up to $1 million per hour of significant outages [6]. Cloud-native architectures contribute to the challenges. Multi-tenancy, scaled-out architecture, distributed query execution, and microservices architecture establish complicated interdependencies that cannot be suitably represented by classical threshold-based or rule-driven systems [7, 8]. According to recent surveys, anomaly detection in cloud environments must be performed with the help of advanced machine learning (ML) and deep learning (DL) models that can process multi-dimensional telemetry data (such as query logs, transaction streams, resource utilization metrics, and network traffic patterns) [9, 10]. New studies indicate the possibility of supervised learning, unsupervised clustering, and ensemble approaches to predictive anomaly detection [11, 12]. Long Short-Term Memory (LSTM) networks, Convolutional Neural Networks (CNNs), and Transformer networks are just a select few types of deep learning networks that have demonstrated specific promise in describing temporal dependencies and intricate patterns in time-series data [13, 14]. At the same time, low-latency data ingestion and inference on data at scale is made possible by stream processing platforms like Apache Kafka and Apache. To solve these problems, this paper presents Resilient Dataflow Intelligence (RDI), a holistic model that combines stream processing, ensemble ML/DL models, and adaptive response. Compared to the current methods, RDI is devised in such a way that it: ● Anticipate anomalies with high accuracy and low latency, before affecting the performance of the system. ● Correlate patterns when distributed in heterogeneous multi-cloud environments. ● Reduce false positives by the adoption of advanced ensemble techniques and adaptive threshold systems. ● Interoperable with the major cloud vendors (AWS, Google Cloud Platform, Microsoft Azure) and deployable across all major cloud providers. Figure 1 Evolution of Database Monitoring from Reactive to Proactive Approaches 1.1. Problem Statement Even with considerable progress in monitoring tools and analytics features, cloud databases still face unpredictable issues such as the loss of query performance, transaction congestion, storage congestion, security breaches, and resource overload [17, 18]. The existing anomaly detection solutions are largely reactive in nature, which poses several critical limitations, and, therefore, the research seeks to resolve these limitations. 1.2. Paper Structure and Objectives In this paper, the researcher aims to accomplish the following three main objectives: To develop and deploy a scalable, vendor-agnostic architecture that incorporates real-time data ingestion, stream processing, machine learning pipelines, and automated response orchestration. 2. To create and test ensemble machine learning systems based on supervised
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 316 learning algorithms, unsupervised algorithms, and deep learning networks. 3. To prove the usefulness of the framework by thoroughly testing it in various cloud platforms and industries. The study contributes to the area of cloud database management in the sense that it is the first to offer a thorough analysis of proactive anomaly detection in multi-cloud setups. We have contributed: (1) a multi-layer RDI framework architecture, (2) real-life exemplification based on three industry case studies demonstrating the outcome of the implementation, (3) a performance evaluation framework of anomaly prediction systems, and (5) practical recommendations on organizations wishing to do the same. 2. Conceptual Review 2.1. Conventional Tracking and Threshold-Based Detection The initial database monitoring solutions used were based on the use of more traditional methods of monitoring, like the usage of a single metric, which included CPU usage, memory usage, disk I/O rates, and query response time, and compared them to a fixed set of predefined limits [1, 19]. The metrics were used to raise alerts to alert operations teams of possible problems when they exceeded the set thresholds. Chen et al. [20] also introduced one of the initial works in performance monitoring of a database, and laid mathematical frameworks on which the threshold could be calculated as per the past performance distributions. Although their method was revolutionary at the moment, it proved to be highly deficient in dynamic cloud settings where workload patterns are highly variable and seasonal. The main weaknesses of threshold-based methods are: (1) they cannot model complex interdependencies between the various metrics, (2) on normal changes in operations, the false positive rates are high, (3) they can not foresee, and (4) the methods are not scaled well in a multi-tenant cloud environment [21]. 2.2. Database Anomaly Detection with the help of machine learning Seeing the weaknesses of rule-based systems, scholars started to consider the use of machine learning to detect anomalies in the database. The initial research was on the supervised learning techniques that had the capability of learning patterns using labeled historical data. Supervised Learning Methodologies: Zhang et al. [22] utilized Support Vector Machines (SVMs) to categorize states of database performance, whereby the accuracy of query performance anomaly was 85% with the use of Support Vector Machines. Their method, however, demanded large amounts of labeled training data and had difficulty with new types of anomalies not in the training data. Rand Forest in Liu and Wang [23] showed a high degree of ability of the database anomaly detectors to noise resistance and the high-dimensional feature space. Their ensemble method lowered the false positive rates by about a quarter of those of single-model options. Unsupervised Learning Methods: Due to the difficulty in acquiring labeled data on anomalies in real-world applications, researchers explored unsupervised techniques that would be able to detect anomalies without being pre-trained on labeled data. Liu et al. [24] proposed Isolation Forest algorithms that were explicitly trained on database telemetry data and reached linear time and high performance in datasets that had different anomaly rates. Their method eventually became popular because of its computational effectiveness and interpretation. 2.3. Time-Series Anomaly Detection based on Deep Learning Deep learning revolutionized the anomaly-detecting capabilities, especially of time-series data, which is prevalent in the monitoring of databases. Recurrent Neural Networks: The use of LSTM networks as a database performance anomaly detector introduced by Malhotra et al. [25] was the first to establish itself as remaining more competitive in long-term temporal dependencies than the conventional statistical techniques. Their method had 92% accuracy in forecasting anomalies in query response time as far as 30 seconds.
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 317 Transformer Architectures: The most recent research by Zhou et al., showed that self-attention mechanisms were better able to learn global dependencies of database telemetry data than recurrent models, providing a new direction of anomaly detection in complex, multi-dimensional data. 3. RDI Framework Architecture and Design 3.1. Architectural Overview RDI has a layered architecture that is scalable, modular, and vendor-independent deployment. The structure is made up of five layers, which are interrelated and collaboratively operate to offer holistic capabilities of anomaly anticipation. Figure 2 Multi-layer Architecture of the Resilient Dataflow Intelligence (RDI) Framework 3.2. Layer-by-Layer Component Description 3.2.1. Data Ingestion Layer RDI is based on the Data Ingestion Layer, which takes in, normalizes, and routes telemetry data of a wide variety of cloud databases. This layer solves the heterogeneity issue, which is a feature of multi-cloud environments. Apache Kafka Integration: RDI uses Apache Kafka as the main streaming platform, which offers fault-tolerant and distributed message queuing, with horizontal scaling. Kafka subjects are divided by identifiers of database instances and allow parallel processing without disrupting the order of messages in the telemetry stream of a particular database. Multi-Cloud Connectors: Custom-developed connectors interface with cloud provider APIs: ● AWS Connector: Support CloudWatch APIs, RDS Performance Insights, and VPC Flow Logs.
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 318 ● Google Cloud Connector: Cloud Monitoring, Cloud SQL Insights, and Stackdriver. ● Azure Connector: Linked to Azure Monitor, SQL Analytics, and Application Insights. 3.2.2. Stream Processing Layer The Stream Processing Layer converts raw telemetry into feature-rich data streams that can be used for machine learning inference. It is built around Apache Flink to enable stateful stream processing with low latency and exact-once guarantees. Real-time Preprocessing: Telemetry data that comes in is processed in a few steps: ● Checks on quality and data validation. ● Forward-fill imputation based on statistical methods of missing value imputation. ● Robust statistical outliers and smoothing. ● Standardization and metric normalization. ● Windowing and Aggregation: ● Time-based windowing algorithms generate feature vectors out of streaming data: ● Tumbling (5 seconds) of real-time measures. ● Sliding windows (30s with 5s slide) for the analysis of the trend. ● Transaction-based analysis session windows. ● Machine Learning Pipeline The Machine Learning Pipeline is using ensemble techniques that can utilize more than one type of algorithm to obtain strong anomaly detection using different workload patterns. Supervised Learning Models: ● Random Forest (RF): A collection of decision trees that are trained using historical data of anomalies that are labeled. ● Gradient Boosting Machines (GBM): Prediction by sequential enhancement in an ensemble. Unsupervised Learning Models: ● Isolation Forest: A tree-based anomaly detection algorithm that is used to isolate anomalies by randomly sampling features. ● DBSCAN Clustering: This is a density-based clustering algorithm that determines the anomalies as those points in sparse areas. Deep Learning Models: ● LSTM Networks: Recurrent neural networks with memory cells created to learn long-term temporal dependencies. ● CNN-LSTM hybrid: This is the model that involves both the use of convolutional layers (which identify local patterns) and LSTM layers (which predict temporal structure). 3.2.3. Anomaly Prediction Engine The Anomaly Prediction Engine aggregates outputs from multiple machine learning models to produce final anomaly predictions with associated confidence scores. Ensemble Fusion Strategies: ● Weighted Voting: They are given weighted models according to their past results on validation data. ● Stacking: A meta-learner is trained to make the best combinations of predictions made by a base model. ● Bayesian Model Averaging Uncertainty quantification based on Bayesian inference. 5. Response Orchestration Layer: Response Orchestration Layer converts predictions of anomalies into responses to be taken, interacting with the cloud provider APIs and enterprise management systems.
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 319 ● Automated Response Actions: Horizontal and vertical scaling activity based on predicted resource constraints: Preventive Scaling. ● Query Optimization: Index Recommendations and automatic query plan modifications. ● Connection pool management: Connection pool sizes are dynamically adjusted in response to the anticipated workload variations 4. Machine Learning Algorithms and Ensemble Strategies 4.1. Choice and Rationales of Algorithms RDI uses a multi-algorithm framework that relies on the fact that the various forms of anomalies have unique features that can be best represented using specialized techniques of detection. The choice of the algorithm covers three criteria: ● Coverage: Capability to identify known patterns of anomalies and new anomalies. ● Temporal Modeling: The ability to record the short-term fluctuations and long-term trends. ● Scalability: Real-time computational efficiency. 4.2. Fusion of Ensemble Strategies and Models. In an ensemble method, predictions are made using various algorithms, and thus, the RDI ensemble is better at prediction and robustness than single models. Weighted Voting Ensemble: the individual models play a role in the final predictions in accordance with their historical performance and the level of confidence they have. Mathematical Formulation: To make an ensemble prediction at time t: P_ensemble(t) = Σ(i=1 to N) w_i(t) × P_i(t) Where: Where P i (t) is the model i's prediction at time t. Where w i(t )= dynamic weight of model i at time t. N is an ensemble of models. ● Stacking Ensemble: A meta-learner is an ensemble of base model predictions assembled through an algorithm that has been trained-usually, usually by logistic regression or neural networks. ● Bayesian Model Averaging: Predictive uncertainty with quantification Probability-based ensemble approach. ● Stacking Ensemble: A meta-learner is an ensemble of base model predictions assembled through an algorithm that has been trained-usually, usually by logistic regression or neural networks. ● Bayesian Model Averaging: Predictive uncertainty with quantification Probability-based ensemble approach. 5. Methodology and Implementation of the Experiment. 5.1. Experimental Environment Set-Up. 5.1.1. Configuration of the Cloud Platform. Amazon Web Services (AWS): ● Database instance: Amazon RDS, MySQL 8.0, and PostgreSQL 13. ● Scalability testing: db.r5.large, db.r5.xlarge, db.r5.2xlarge instances. ● Storage: Provisioned IOPS GP2 SSD with consistency in performance. ● Monitoring: CloudWatch metrics (1-minute resolution).
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 320 Google Cloud Platform (GCP): ● Database service: MySQL 8.0 and PostgreSQL 13 Cloud SQL. ● Types of machine: db-n1-standard-2, db-n1-standard-4, db-n1-highmem-4. ● Storage: This represents a type of persistent disk, which automatically backs up and therefore can be referred to as an SSD persistent disk. ● Monitoring: Cloud Monitoring with Programmable metric gathering. Microsoft Azure: ● Database service: Azure SQL Database, as well as the Azure Database for PostgreSQL. ● General Purpose and Business Critical configurations: Service tiers. ● Compute sizes 2-8 vCores of different memory allocations. ● Monitoring: Application Insights integration with Azure Monitor. 5.1.2. Benchmarks and Generation of Workloads. TPC-C Benchmark ( E-commerce Simulation): ● •Workload features: The high-frequency transactional processing. ● •Parallel connections: 100-1000 parallel connections. ● •Mix of transactions: New orders (45%), Payments (43%), Order status (4%). ● 50-100GB data (20-100 kb warehouses 50-100 kb database) ● |human|>50-100GB data (20-100 kb warehouses 50-100 kb database) TPC-H Benchmark (Analytics Simulation): ● Features of the workload: Analytical queries that return very large result sets. ● Query complexity: 22 standardized analytical queries of varying complexity. ● Scale / Data size: Three benchmark scales — 10 GB, 100 GB, and 1 TB datasets (sometimes referred to as scale factors 10, 100, and 1,000). MIMIC-IV Healthcare Dataset: Characteristics of datasets: De-identified electronic health records. Storage: 40GB(compressed) 200GB(uncompressed) Patterns: Query patterns: Patient lookups, clinical decision support, and Research analytics. 6. Findings and Analysis of Performance. 6.1. General Performance Report. RDI demonstrated higher performance in all the metrics evaluated as opposed to the baseline strategies: Key Performance Indicators: ● Prediction Accuracy: The overall workload mean accuracy is 93.2 percent. ● F1Score: The mean F1-score is 0.91, which represents a balanced F1-score. ● False Positive: 2.1% average false positive rate (reduction of 68 percent compared to baselines) ● Early Detection: 87% of the anomalies with over 30 seconds of lead time are observed. ● Inference Latency: End-to-end average prediction latency of 12.8ms. ● Downtime Reduction: 48 percent average reduction of system downtime. ● Cost Saving: 31% reduction of operational costs on average. 6.2. Performances by Type of Workload. 6.2.1. E-commerce TPC-C Results. Experimental Configuration:
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 321 ● Platform AWS RDS MySQL 8.0 on db.r5.xlarge instances. ● Workload: 500 concurrent users, 15000 TPS average. ● Execution: 72-hour continuous execution. ● Abnormalities: 1,200 injected abnormalities of 6 types. Table 1 E-commerce Workload Performance Comparison Metric RDI Threshold-Based Single ML Deep Learning Commercial Accuracy (%) 94.1 71.2 84.3 89.7 82.5 Precision (%) 92.8 68.9 81.5 87.2 79.8 Recall (%) 95.6 74.1 87.8 92.1 85.3 F1-Score 0.94 0.71 0.84 0.89 0.82 False Positive Rate (%) 1.8 12.4 4.7 3.2 6.1 Mean Lead Time (seconds) 42.3 N/A 28.7 35.9 19.2 Inference Latency (ms) 11.2 3.1 45.8 67.4 28.9 6.2.2. Cross-Platform Consistency Analysis RDI demonstrated consistent performance across all three major cloud platforms: Table 2 Cross-Platform Performance Consistency Platform Accuracy (%) F1-Score Latency (ms) Throughput (events/sec) AWS RDS 93.8 ± 1.2 0.92 ± 0.02 12.1 ± 2.3 98,400 ± 3,200 GCP Cloud SQL 92.9 ± 1.4 0.91 ± 0.03 13.2 ± 2.7 95,800 ± 4,100 Azure SQL DB 92.7 ± 1.6 0.90 ± 0.03 13.8 ± 2.9 94,200 ± 3,800 Overall Average 93.2 ± 1.4 0.91 ± 0.03 13.0 ± 2.6 96,100 ± 3,700
Global Journal of Engineering and Technology Advances, 2025, 24(03), 314-327 322 Figure 3 Performance Comparison - Traditional vs RDI-Enabled Systems 6.3. Ensemble Method Effectiveness Analysis of individual algorithm contributions within the RDI ensemble: Table 3 Individual Algorithm vs. Ensemble Performance Algorithm Accuracy (%) Precision (%) Recall (%) F1-Score Latency (ms) Random Forest 87.3 84.9 89.8 0.87 8.4 XGBoost 89.1 87.2 91.2 0.89 12.7 Isolation Forest 82.4 78.6 86.9 0.83 6.2 DBSCAN 79.8 76.1 84.2 0.80 18.9 LSTM 88.7 86.4 91.3 0.89 45.3 CNN-LSTM 87.9 85.7 90.4 0.88 52.1 RDI Ensemble 93.2 91.8 94.7 0.93 13.0 The ensemble approach provides 4.1-13.4% accuracy improvement over individual algorithms while maintaining reasonable latency.